07
Attention stage

Query · Key · Value

Attention is computed with three projections of each token, each asking a different question.

The Attention chapter showed what queries, keys, and values do. This one shows where they come from: every token's embedding is projected through three separate learned weight matrices, producing three separate vectors for that one token.

Embedding for “it”
× W^Q
Query
× W^K
Key
× W^V
Value

Same embedding, three separate learned weight matrices, three differently-shaped output vectors. The bars are illustrative, not a real trained model's weights.

qi=xiWQki=xiWKvi=xiWVq_i = x_i W^Q \qquad k_i = x_i W^K \qquad v_i = x_i W^V

xix_i is the position-encoded embedding for token ii. WQW^Q, WKW^K, and WVW^V are ordinary learned weight matrices, no different in kind from any other parameter in the model. Nothing about them is fixed or hand-designed; training is what shapes them into projections that make attention useful.

Every token gets its own query, key, and value, including the query token itself. That's why a token can, and often does, attend partly to its own position.

Why not compare embeddings directly?

The simplest scheme would score two tokens by the dot product of their vectors, xi⋅xjx_i \cdot x_j. It has two problems. First, it is symmetric: xi⋅xjx_i \cdot x_j always equals xj⋅xix_j \cdot x_i, but relationships between words have a direction. A pronoun goes looking for a noun; the noun isn't looking for the pronoun in the same way. Second, one vector would have to do three jobs at once: describe what a token is looking for, advertise what it contains, and supply the content it hands over. Those goals pull in different directions. Three separate projections give each job its own view of the token.

Substituting qi=xiWQq_i = x_i W^Q and kj=xjWKk_j = x_j W^K shows why this fixes the symmetry:

qi⋅kj=xi WQ(WK)T xjTq_i \cdot k_j = x_i \, W^Q (W^K)^{T} \, x_j^{T}

The matrix WQ(WK)TW^Q (W^K)^T sits between the two token vectors, and in general it is not symmetric, so the score from ii to jj can differ from the score from jj to ii.

Shapes and sizes

WQW^Q and WKW^K have shape d×dkd \times d_k, and WVW^V has shape d×dvd \times d_v. In the demo above d=4d = 4 and each output has 4 numbers, so you can read them. In GPT-3, d=12,288d = 12{,}288 and each attention head projects down to 128 numbers (96 heads of 128 add back up to 12,288, as the next chapter shows). Counting all heads together, the four matrices of an attention layer (WQW^Q, WKW^K, WVW^V and the output matrix WOW^O) hold about 4d24d^2 parameters: roughly 2.4 million per layer in GPT-2 small.

What training does to these matrices

Nobody designs WQW^Q, WKW^K or WVW^V. They start random and are adjusted by gradient descent (chapter 16), because whatever they do affects how well the next token gets predicted. If letting pronoun queries match noun keys lowers the loss, gradients move the matrices in that direction. The labels "asking", "advertising" and "delivering" are our interpretation of what trained matrices tend to end up doing. The mathematics is only three matrix multiplications.

A three-token example by hand

Take a tiny head with dk=2d_k = 2, so dk≈1.41\sqrt{d_k} \approx 1.41. The query for "it" is (1.0, 0.5)(1.0,\,0.5), and three keys are on offer:

Scores and weights for the query "it"

key        k             q·k    ÷√d_k    e^score   weight
robot     (2.0, 1.0)    2.50    1.77      5.86     0.71
patient   (0.5, -1.0)   0.00    0.00      1.00     0.12
it        (0.0, 1.0)    0.50    0.35      1.42     0.17
                                          -----    ----
                                          8.28     1.00

The output is 0.71 vrobot+0.12 vpatient+0.17 vit0.71\,v_{\text{robot}} + 0.12\,v_{\text{patient}} + 0.17\,v_{\text{it}}. The Simulator page repeats this arithmetic one operation at a time, with the lab's stand-in vectors.

So far there is one set of three matrices, and therefore one attention pattern per layer. The next chapter asks what happens if the model runs several sets side by side.

Key Takeaway

Query, key, and value are not three different kinds of information floating around. They're the same embedding, passed through three separate learned projections, each shaped by training to serve a different role in the attention computation.