13
Probability stage

Softmax

Softmax turns raw scores into a probability distribution that sums to one.

Raw scores→Exponential→Normalize→Probabilities

Softmax does two things to a list of logits: it exponentiates every value (which makes everything positive, and makes larger scores disproportionately larger), then divides each one by the total, so the whole list sums to exactly 1.

softmax(xi)=exi∑jexj\text{softmax}(x_i) = \dfrac{e^{x_i}}{\sum_j e^{x_j}}
1.00
0.1 (sharper)1.02.0 (flatter)
moving
47.5%
running
23.8%
falling
16.3%
walking
8.6%
heavy
2.5%
silent
0.9%
blue
0.3%
yesterday
0.1%

The bars always sum to 100% at every temperature. That's the point of normalizing; temperature just controls how concentrated the distribution is beforehand.

Temperature isn't part of softmax's definition. It's a knob applied before softmax, dividing every logit by it first. A temperature below 1 sharpens the distribution toward the top candidate; above 1 flattens it toward uniform. At the low extreme it behaves almost like always picking the top score; at the high extreme, almost like picking uniformly at random.

The formula, one step at a time

P(yi)=ezi∑j=1∣V∣ezjP(y_i) = \dfrac{e^{z_i}}{\sum_{j=1}^{|V|} e^{z_j}}

There are two steps. Exponentiating each logit makes every score positive and keeps the order (a higher logit still gives a higher value), and it stretches gaps, because eze^{z} grows faster the larger zz is. Dividing by the total then forces the values to add up to 1, so they form a probability distribution. The name fits: it is a soft version of picking the maximum. The top score gets the most probability, but the others keep some.

Why the exponential and not, say, squaring? Squaring would rank -3 above 1. The exponential works for any real score, it is smooth so gradients exist everywhere, and it turns score differences into probability ratios, the property from the previous chapter. It also pairs neatly with the training loss: with cross-entropy, the gradient with respect to a logit is simply pi−yip_i - y_i (chapter 16).

A shift changes nothing

Adding the same number cc to every logit leaves the output unchanged, because ezi+c/∑jezj+ce^{z_i + c} / \sum_j e^{z_j + c} cancels the common factor ece^c. Implementations use this to subtract the largest logit before exponentiating, which prevents overflow on large scores without altering the answer. It also means the absolute level of the logits carries no information. Only their order and gaps do.

Temperature

PT(yi)=ezi/T∑jezj/TP_T(y_i) = \dfrac{e^{z_i / T}}{\sum_j e^{z_j / T}}

Temperature TT divides the logits before softmax. With T<1T < 1 the gaps widen and probability concentrates on the leaders. With T>1T > 1 the gaps shrink and probability spreads out. As T→0T \to 0 the distribution collapses onto the top token, and as T→∞T \to \infty it approaches uniform. For the logits in this lab:

The same eight logits at three temperatures

token       T = 0.5    T = 1    T = 2
moving        71.2%    47.5%    31.2%
running       17.9%    23.8%    22.1%
falling        8.4%    16.3%    18.3%
walking        2.3%     8.6%    13.3%
heavy          0.2%     2.5%     7.1%
silent         0.0%     0.9%     4.4%
blue           0.0%     0.3%     2.3%
yesterday      0.0%     0.1%     1.4%

Temperature is a setting you choose when generating, not something the model learned, and it changes nothing about the logits themselves.

Two softmaxes in one model

The same function appears twice in a transformer. Inside every attention head it turns a row of scores into mixing weights over positions. Here, at the very end, it turns logits into a distribution over the vocabulary. The formula is identical and the jobs differ: the first decides where to read from, the second decides what to say next.

The output is the model's own estimate of P(xt∣x<t)P(x_t \mid x_{<t}). It is a valid distribution, but nothing guarantees it matches real-world frequencies, and models tuned to follow instructions tend to be less well calibrated than base models. A distribution still doesn't choose anything. Turning it into one token is the job of sampling, next.

Key Takeaway

Softmax turns arbitrary real-valued scores into a valid probability distribution, always positive and always summing to 1, and temperature controls how peaked or flat that distribution ends up.