15
Probability stage

Generation

The chosen token is appended to the context, and the whole process runs again.

Every chapter so far has covered one forward pass: one full trip from text to a chosen next token. Generating more than one token just means doing that same trip again, with the new token folded into the context each time. Step through it below.

“The robot is…”

1. Context2. Transformer3. Logits4. Sample5. Append

The current context (everything generated so far, plus the original prompt) is fed in as-is.

This loop is why a model can only "see" what fits inside its context window. Mathematically, every new token depends on the whole context, so each step could recompute everything from the start. Real systems avoid that: earlier tokens' keys and values can never change (causal masking sees to that), so they are kept in a cache and only the newest token is computed at each step. The cache saves work without adding information. Either way, the weights are identical at every step, and nothing is carried between steps except the growing context.

One pass gives one token

A full forward pass (tokens, embeddings, position, every block, the final projection, softmax) produces a single thing: a distribution over the next token. To get a whole reply, the pass is run again and again, each time with one more token in the context. This kind of generation is called autoregressive: the model's own earlier outputs become its later inputs.

The loop as a formula

P(x1,…,xT)=∏t=1TP(xt∣x<t)P(x_1, \dots, x_T) = \prod_{t=1}^{T} P\big(x_t \mid x_{<t}\big)

This is the chain rule of probability, not an approximation: any distribution over sequences can be split into one conditional distribution per position. The model only ever learns the factors on the right, each one a next-token distribution. Generating a text means drawing those factors one after another, and the probability of the whole text is their product.

The loop in pseudocode

context = prompt_tokens
while not finished:
    logits = model(context)[-1]      # last position only
    probs  = softmax(logits / T)     # temperature
    token  = sample(probs)           # top-k / top-p
    context.append(token)
    finished = token == END or len(context) >= limit

Why it is sequential

Each token depends on the ones chosen before it, random choices included, so the future can't be computed ahead of time. During training this problem doesn't arise, because the true text is already known and every position can be predicted in parallel (chapter 16). At generation time that shortcut is gone, which is why writing a long answer takes noticeably longer than reading a long prompt. A request has two phases: first the prompt is processed, all its tokens in parallel; then the reply is produced one token per pass.

The KV cache

Redoing every step from scratch would repeat almost all of the work, since the keys and values of earlier tokens haven't changed. Causal masking guarantees that they can't: a token never attends to anything later than itself, so appending a new token alters nothing computed for the older ones. Real systems therefore keep each layer's keys and values for every token so far (the KV cache). At each step only the new token goes through the network. Its query attends over the cached keys and values, and its own key and value are added to the cache. The result is identical to recomputing everything, only faster.

The cost moves to memory. For each token the cache holds a key and a value in every layer. For a GPT-3-sized model that is 2⋅96⋅12,2882 \cdot 96 \cdot 12{,}288 numbers, roughly 4.7 MB per token at 2 bytes each, so a 2,048-token conversation carries nearly 10 GB of cache. That is a large part of why long contexts are expensive to serve, and why techniques such as grouped-query attention (chapter 8) exist to shrink it.

What the loop can't do

Once a token is emitted it stays. The model can't go back and revise an earlier word. It can only continue from what it wrote, and a mistake becomes part of the context that shapes everything after it. What the network computes internally may well anticipate later words, but it commits to one token at a time. The model also has no memory beyond its context window: what isn't in the current context, it can't use. A chat assistant that seems to remember earlier turns is being sent them again each time.

Everything in this chapter happens with the weights frozen. Nothing the model reads or writes during generation changes it. How the weights got their values is the subject of the next chapter.

Key Takeaway

Autoregressive generation is one forward pass, repeated: context in, one new token out, append, repeat. Everything from Text through Sampling happens again at every single step.