Logits
The final layer produces one raw, unnormalized score per possible next token.
After the last transformer block, every position in the sequence has a final vector. For predicting the next token, only the vector at the last position matters. It gets multiplied by one more weight matrix (the "unembedding") to produce one score per token in the entire vocabulary. For context "The robot is", here are the top candidates:
These are raw scores for "The robot is ___". Notice they can be negative, and there's no reason they'd sum to any particular number. That's what makes them logits and not probabilities.
A vocabulary has anywhere from tens of thousands to a few hundred thousand entries, so this is really one score for every single one of them; the chart above only shows the highest ones. Everything else got a logit too, usually far below the leaders. Only the differences between logits matter, not whether they are positive or negative.
From one vector to one score per token
The vector at the last position, , has width and holds everything the model has worked out about what comes next (in a pre-norm model it is the output of the final LayerNorm). To compare it with every candidate token, the model multiplies it by the unembedding matrix , which has one column for each token in the vocabulary:
Entry of is , where is column of : the same kind of dot product as an attention score. Each token has an output vector , and its score is large when points in the same direction. Many models, GPT-2 among them, tie the two tables so that . The vector that represents a token on the way in is then reused to score it on the way out. Many recent models leave out the bias .
Which positions get scored
During generation only the last position is projected, since only its prediction is needed. During training every position is projected at once: position , which sees only tokens up to thanks to the causal mask, predicts token , so one pass over tokens yields predictions to learn from (chapter 16).
What a logit means
Logits are unnormalized. They can be any real number, and a negative logit is not a negative probability. What matters is how they compare. Two facts follow from the softmax in the next chapter. Adding the same constant to every logit changes nothing, and the gap between two logits is the logarithm of how much more likely one token is than the other:
In the chart, "moving" (4.81) and "running" (4.12) are 0.69 apart, and , so "moving" is about twice as likely as "running" whatever else is in the vocabulary. That relationship is only easy to see once the scores are normalized, which is why the chart shows no percentages yet.
The cost of a large vocabulary
This projection is one of the biggest single matrices in a model: weights, about 38.6 million in GPT-2 small and over 500 million for a 128,000-token vocabulary at width 4,096. It is also computed at every generation step. Vocabulary size is therefore a trade-off: a bigger vocabulary needs fewer tokens per text (chapter 3) but makes this layer, and the embedding table, larger.
Nothing has been chosen yet. There is one raw score per token, and they add up to nothing in particular. The next chapter turns them into probabilities.