Sampling
A next token is drawn from the distribution, not always the single most likely one.
A probability distribution alone doesn't pick a token; something has to draw from it. A few common strategies for doing that:
- Greedy: always take the single highest-probability token. Deterministic, and often repetitive over a long generation.
- Temperature: reshape the whole distribution (see the previous chapter), then sample from all of it.
- Top-k: keep only the k highest-probability tokens, renormalize, then sample.
- Top-p (nucleus): keep the smallest set of top tokens whose probabilities add up to at least p, renormalize, then sample.
Dimmed tokens are excluded from this draw. Greedy always excludes everything but the top token, so every draw gives the same result. The other strategies leave room for variation.
Top-k uses a fixed number of candidates regardless of how confident the distribution is; top-p adapts instead: a very confident distribution might keep just one or two tokens, while a very uncertain one might keep a dozen. Neither is strictly "better"; they're different ways of trading off coherence against variety.
The model computes a distribution; you choose how to pick from it
Everything up to softmax belongs to the network. Picking a token from the distribution is a separate step, called decoding, that lives outside it. The same model with different decoding settings writes noticeably different text, which is also why one prompt can give different answers from one run to the next.
Greedy, and the long tail
Always taking the top token is deterministic but rarely best. The most likely token at each step doesn't add up to the most likely text, and long greedy outputs tend to fall into repetitive loops. Sampling in proportion to the probabilities brings variety, and with it a different problem. The demo's distribution covers eight candidates so it fits on screen. A real softmax covers the whole vocabulary, where tens of thousands of unlikely tokens each hold a tiny probability. Individually they are negligible, but they add up: 50,000 tokens at 0.0002% each hold 10% of the probability. Drawing from the untouched distribution would therefore pick a poor token about one time in ten. Truncation methods exist to cut that tail off.
Top-k and top-p
Top-k keeps the most probable tokens and renormalizes. Top-p, also called nucleus sampling, keeps the smallest set of tokens whose probabilities add up to at least , so the number kept adapts to how confident the model is. With the demo's logits at and , the running totals go 47.5%, 71.4%, 87.7%, so the first three tokens are kept. Renormalizing over that set :
Top-k with would keep the same three tokens here. They part ways when the distribution changes. If one token holds 95% of the probability, top-p keeps just that one, while top-k still lets two long shots through. If the model is torn between twenty plausible words, top-p keeps many of them and top-k cuts most away.
How the settings combine
A common order is this: temperature is applied to the logits, top-k and top-p filter the result, the survivors are renormalized, and one token is drawn. Libraries differ in details, so check the documentation of the one you use. Other tools exist. Min-p keeps tokens whose probability is at least some fraction of the top token's. Penalties can discourage repeating recent tokens. Beam search, used mostly in translation and rarely in chat, tracks several candidate sequences rather than committing to one token at a time.
When does it stop?
Generation ends when the model produces a special end-of-sequence token or reaches a length limit. That token is just another entry in the vocabulary, one the model learned to emit when a response is complete.
The chosen token still has to go somewhere. The next chapter feeds it back into the context and starts the whole process over.