Training
Parameters are nudged, over and over, so predicted tokens match real ones more often.
Every earlier chapter described a model that already knows how to do this. Training is how it gets there: the same predict-a-next-token computation, run over and over, with parameters nudged a little after each example so the prediction gets a little better.
Input: "The cat" → target: "sat"
This is a real (if tiny) gradient descent loop: four learnable numbers, one example, updated with the actual softmax cross-entropy gradient each step. A production model does the same update rule across billions of parameters and millions of examples, not four numbers and one sentence.
The loss only cares about the probability assigned to the correct token: the lower that probability, the higher the loss. The gradient of this loss with respect to each candidate's score works out to something clean: subtract 1 from the correct candidate's probability, leave the rest as their probability. That's exactly the adjustment the demo above applies at every step.
This is a genuinely tiny model: four numbers, one example, no embeddings or attention involved at all. A real model repeats the same underlying update rule, gradient descent, across billions of parameters and enormous datasets, using backpropagation to compute every parameter's gradient at once. The mechanism is the same; the scale is not.
Same network, different job
Training and inference side by side
training inference weights updated frozen text known, from a corpus your prompt next token known: the target unknown: sampled positions all predicted at once one new token per step
Inference, which the earlier chapters described, uses the network. Training is the process that set the numbers inside it. Both run the same forward computation. What differs is what happens around it.
The objective: predict the next token
Take a stretch of real text, . At each position the model outputs a distribution over the next token, and the loss measures how much probability it gave to the token that actually came next:
This is the cross-entropy loss. If the model gave the right token probability 1, that term is 0. At probability 0.1 it is , and at 0.01 about 4.6, so confident mistakes are punished hard. Minimizing the average is the same as maximizing the probability the model assigns to the training text, by the chain rule from the previous chapter. Its exponential, , is called perplexity: a loss of 2.3 is like being as unsure at each step as a fair ten-sided die.
All positions at once
Because of the causal mask, position can't see token or anything after it, so a single pass over a -token passage gives separate predictions, each scored against the true next token. The true earlier tokens are fed in, not the model's own guesses, a setup called teacher forcing. That is why training is so much more parallel than generation.
How the weights change
The forward pass computes the loss. Backpropagation then applies the chain rule backward through every layer to get the gradient : for every weight in the model, how much the loss would change if that weight changed a little. At the output the gradient takes a clean form. For softmax with cross-entropy, the gradient with respect to logit is:
where is 1 for the true token and 0 for every other. Tokens that received probability they shouldn't have are pushed down in proportion to it, and the right token is pushed up by exactly the probability it was missing. Then the update moves each weight a small step against its gradient:
Here is the learning rate. Production training uses a refinement called Adam, which scales each step using running averages of past gradients, and averages the gradient over a batch of many sequences so one odd example doesn't dominate. The loop repeats over hundreds of billions to trillions of tokens. The residual connections from chapter 10 are what keep the gradient alive through dozens of layers.
What the demo shows
In the demo the four logits are themselves the trainable numbers, and each step uses the real gradient above with a plain update. Two things are worth watching. Each step moves probability toward the target. And the loss drops fast while the model is most wrong, then slows as it gets close, because shrinks as the target's probability approaches 1.
What training produces
Nobody writes grammar rules or facts into the model. Whatever helps predict the next token gets reinforced, and the model ends up with a huge set of numbers that encode regularities of its training text, including a great deal of what we would call grammar, facts and style. Training is the slow, expensive stage, done once per model. Everything in the earlier chapters is the cheap part that runs each time you use the result.
Pretraining on raw text produces a model that continues text well. Turning it into an assistant is the subject of the next chapter.