14
Probability stage

Sampling

A next token is drawn from the distribution, not always the single most likely one.

A probability distribution alone doesn't pick a token; something has to draw from it. A few common strategies for doing that:

  • Greedy: always take the single highest-probability token. Deterministic, and often repetitive over a long generation.
  • Temperature: reshape the whole distribution (see the previous chapter), then sample from all of it.
  • Top-k: keep only the k highest-probability tokens, renormalize, then sample.
  • Top-p (nucleus): keep the smallest set of top tokens whose probabilities add up to at least p, renormalize, then sample.
Temperature0.80
moving
running
falling
walking
heavy
silent
blue
yesterday

Dimmed tokens are excluded from this draw. Greedy always excludes everything but the top token, so every draw gives the same result. The other strategies leave room for variation.

Top-k uses a fixed number of candidates regardless of how confident the distribution is; top-p adapts instead: a very confident distribution might keep just one or two tokens, while a very uncertain one might keep a dozen. Neither is strictly "better"; they're different ways of trading off coherence against variety.

The model computes a distribution; you choose how to pick from it

Everything up to softmax belongs to the network. Picking a token from the distribution is a separate step, called decoding, that lives outside it. The same model with different decoding settings writes noticeably different text, which is also why one prompt can give different answers from one run to the next.

Greedy, and the long tail

Always taking the top token is deterministic but rarely best. The most likely token at each step doesn't add up to the most likely text, and long greedy outputs tend to fall into repetitive loops. Sampling in proportion to the probabilities brings variety, and with it a different problem. The demo's distribution covers eight candidates so it fits on screen. A real softmax covers the whole vocabulary, where tens of thousands of unlikely tokens each hold a tiny probability. Individually they are negligible, but they add up: 50,000 tokens at 0.0002% each hold 10% of the probability. Drawing from the untouched distribution would therefore pick a poor token about one time in ten. Truncation methods exist to cut that tail off.

Top-k and top-p

Top-k keeps the kk most probable tokens and renormalizes. Top-p, also called nucleus sampling, keeps the smallest set of tokens whose probabilities add up to at least pp, so the number kept adapts to how confident the model is. With the demo's logits at T=1T = 1 and p=0.8p = 0.8, the running totals go 47.5%, 71.4%, 87.7%, so the first three tokens are kept. Renormalizing over that set SS:

P′(i)=P(i)∑j∈SP(j),0.4750.877≈54%,0.2380.877≈27%,0.1630.877≈19%P'(i) = \dfrac{P(i)}{\sum_{j \in S} P(j)}, \qquad \dfrac{0.475}{0.877} \approx 54\%, \quad \dfrac{0.238}{0.877} \approx 27\%, \quad \dfrac{0.163}{0.877} \approx 19\%

Top-k with k=3k = 3 would keep the same three tokens here. They part ways when the distribution changes. If one token holds 95% of the probability, top-p keeps just that one, while top-k still lets two long shots through. If the model is torn between twenty plausible words, top-p keeps many of them and top-k cuts most away.

How the settings combine

A common order is this: temperature is applied to the logits, top-k and top-p filter the result, the survivors are renormalized, and one token is drawn. Libraries differ in details, so check the documentation of the one you use. Other tools exist. Min-p keeps tokens whose probability is at least some fraction of the top token's. Penalties can discourage repeating recent tokens. Beam search, used mostly in translation and rarely in chat, tracks several candidate sequences rather than committing to one token at a time.

When does it stop?

Generation ends when the model produces a special end-of-sequence token or reaches a length limit. That token is just another entry in the vocabulary, one the model learned to emit when a response is complete.

The chosen token still has to go somewhere. The next chapter feeds it back into the context and starts the whole process over.

Key Takeaway

Sampling is the step that turns a probability distribution into an actual choice. Which strategy is used changes how repetitive or how varied generated text tends to be; it doesn't change how the distribution itself was computed.