16
Training stage

Training

Parameters are nudged, over and over, so predicted tokens match real ones more often.

Every earlier chapter described a model that already knows how to do this. Training is how it gets there: the same predict-a-next-token computation, run over and over, with parameters nudged a little after each example so the prediction gets a little better.

Predict→Loss→Gradient→Update weights→ repeat

Input: "The cat" → target: "sat"

sat
25.0%
ran
25.0%
slept
25.0%
jumped
25.0%
Loss = −log P("sat")1.386
Learning rate0.30
step 0

This is a real (if tiny) gradient descent loop: four learnable numbers, one example, updated with the actual softmax cross-entropy gradient each step. A production model does the same update rule across billions of parameters and millions of examples, not four numbers and one sentence.

Cross-entropy loss for one example
L=−log⁡P(y)L = -\log P(y)

The loss only cares about the probability assigned to the correct token: the lower that probability, the higher the loss. The gradient of this loss with respect to each candidate's score works out to something clean: subtract 1 from the correct candidate's probability, leave the rest as their probability. That's exactly the adjustment the demo above applies at every step.

This is a genuinely tiny model: four numbers, one example, no embeddings or attention involved at all. A real model repeats the same underlying update rule, gradient descent, across billions of parameters and enormous datasets, using backpropagation to compute every parameter's gradient at once. The mechanism is the same; the scale is not.

Same network, different job

Training and inference side by side

                 training              inference
weights          updated               frozen
text             known, from a corpus  your prompt
next token       known: the target     unknown: sampled
positions        all predicted at once one new token per step

Inference, which the earlier chapters described, uses the network. Training is the process that set the numbers inside it. Both run the same forward computation. What differs is what happens around it.

The objective: predict the next token

Take a stretch of real text, x1,…,xTx_1, \dots, x_T. At each position the model outputs a distribution over the next token, and the loss measures how much probability it gave to the token that actually came next:

L(θ)=−1T∑t=1Tlog⁡Pθ(xt∣x<t)\mathcal{L}(\theta) = -\dfrac{1}{T}\sum_{t=1}^{T} \log P_\theta\big(x_t \mid x_{<t}\big)

This is the cross-entropy loss. If the model gave the right token probability 1, that term is 0. At probability 0.1 it is −ln⁡0.1≈2.3-\ln 0.1 \approx 2.3, and at 0.01 about 4.6, so confident mistakes are punished hard. Minimizing the average is the same as maximizing the probability the model assigns to the training text, by the chain rule from the previous chapter. Its exponential, eLe^{\mathcal{L}}, is called perplexity: a loss of 2.3 is like being as unsure at each step as a fair ten-sided die.

All positions at once

Because of the causal mask, position tt can't see token tt or anything after it, so a single pass over a TT-token passage gives TT separate predictions, each scored against the true next token. The true earlier tokens are fed in, not the model's own guesses, a setup called teacher forcing. That is why training is so much more parallel than generation.

How the weights change

The forward pass computes the loss. Backpropagation then applies the chain rule backward through every layer to get the gradient ∇θL\nabla_\theta \mathcal{L}: for every weight in the model, how much the loss would change if that weight changed a little. At the output the gradient takes a clean form. For softmax with cross-entropy, the gradient with respect to logit ii is:

∂L∂zi=pi−yi\dfrac{\partial \mathcal{L}}{\partial z_i} = p_i - y_i

where yiy_i is 1 for the true token and 0 for every other. Tokens that received probability they shouldn't have are pushed down in proportion to it, and the right token is pushed up by exactly the probability it was missing. Then the update moves each weight a small step against its gradient:

θ←θ−η ∇θL\theta \leftarrow \theta - \eta\,\nabla_\theta \mathcal{L}

Here η\eta is the learning rate. Production training uses a refinement called Adam, which scales each step using running averages of past gradients, and averages the gradient over a batch of many sequences so one odd example doesn't dominate. The loop repeats over hundreds of billions to trillions of tokens. The residual connections from chapter 10 are what keep the gradient alive through dozens of layers.

What the demo shows

In the demo the four logits are themselves the trainable numbers, and each step uses the real gradient above with a plain update. Two things are worth watching. Each step moves probability toward the target. And the loss drops fast while the model is most wrong, then slows as it gets close, because pi−yip_i - y_i shrinks as the target's probability approaches 1.

What training produces

Nobody writes grammar rules or facts into the model. Whatever helps predict the next token gets reinforced, and the model ends up with a huge set of numbers that encode regularities of its training text, including a great deal of what we would call grammar, facts and style. Training is the slow, expensive stage, done once per model. Everything in the earlier chapters is the cheap part that runs each time you use the result.

Pretraining on raw text produces a model that continues text well. Turning it into an assistant is the subject of the next chapter.

Key Takeaway

Training repeats one loop: predict, measure how wrong the prediction was with a loss function, compute how each parameter should change to reduce that loss, and update every parameter slightly. Repetition and scale are what turn this into a capable model.