Query · Key · Value
Attention is computed with three projections of each token, each asking a different question.
The Attention chapter showed what queries, keys, and values do. This one shows where they come from: every token's embedding is projected through three separate learned weight matrices, producing three separate vectors for that one token.
Same embedding, three separate learned weight matrices, three differently-shaped output vectors. The bars are illustrative, not a real trained model's weights.
is the position-encoded embedding for token . , , and are ordinary learned weight matrices, no different in kind from any other parameter in the model. Nothing about them is fixed or hand-designed; training is what shapes them into projections that make attention useful.
Every token gets its own query, key, and value, including the query token itself. That's why a token can, and often does, attend partly to its own position.
Why not compare embeddings directly?
The simplest scheme would score two tokens by the dot product of their vectors, . It has two problems. First, it is symmetric: always equals , but relationships between words have a direction. A pronoun goes looking for a noun; the noun isn't looking for the pronoun in the same way. Second, one vector would have to do three jobs at once: describe what a token is looking for, advertise what it contains, and supply the content it hands over. Those goals pull in different directions. Three separate projections give each job its own view of the token.
Substituting and shows why this fixes the symmetry:
The matrix sits between the two token vectors, and in general it is not symmetric, so the score from to can differ from the score from to .
Shapes and sizes
and have shape , and has shape . In the demo above and each output has 4 numbers, so you can read them. In GPT-3, and each attention head projects down to 128 numbers (96 heads of 128 add back up to 12,288, as the next chapter shows). Counting all heads together, the four matrices of an attention layer (, , and the output matrix ) hold about parameters: roughly 2.4 million per layer in GPT-2 small.
What training does to these matrices
Nobody designs , or . They start random and are adjusted by gradient descent (chapter 16), because whatever they do affects how well the next token gets predicted. If letting pronoun queries match noun keys lowers the loss, gradients move the matrices in that direction. The labels "asking", "advertising" and "delivering" are our interpretation of what trained matrices tend to end up doing. The mathematics is only three matrix multiplications.
A three-token example by hand
Take a tiny head with , so . The query for "it" is , and three keys are on offer:
Scores and weights for the query "it"
key k q·k ÷√d_k e^score weight
robot (2.0, 1.0) 2.50 1.77 5.86 0.71
patient (0.5, -1.0) 0.00 0.00 1.00 0.12
it (0.0, 1.0) 0.50 0.35 1.42 0.17
----- ----
8.28 1.00The output is . The Simulator page repeats this arithmetic one operation at a time, with the lab's stand-in vectors.
So far there is one set of three matrices, and therefore one attention pattern per layer. The next chapter asks what happens if the model runs several sets side by side.