Attention
Each token gathers information from the other tokens around it, weighting each by how relevant it is.
By now each token has a vector that says which token it is and where it sits, and those vectors have never interacted. Yet meaning depends on other words: "bank" means something different next to "river" than next to "loan", and the "it" in the sentence below points back to someone. Attention is the operation that lets each token's vector be updated with information from other tokens. It is the only place in a transformer block where information moves between tokens.
Example sentence: “The robot carried the patient because it was tired.” It’s deliberately ambiguous: who does "it" refer to?
Token inspector
- Query (row)
- it
- Key (column)
- robot
- Raw score (QKᵀ/√d)
- 1.90
- Attention weight
- 0.629
Click any cell. Row it's weights always sum to 1 across every column. That's exactly what a row of softmax output means.
These weights are hand-designed for this example, not the output of a trained model. Their only purpose is to make the shape of the computation easy to see. A real model learns attention patterns like these through training, never by hand.
Notice that “it” attends most strongly back to “robot,” not to “patient.” The sentence itself leaves that relationship ambiguous on the surface, yet attention can still surface a pattern like this, even though nothing in the sentence states it explicitly.
Intuitive
Attention is a mechanism that lets a token's representation take into account information from other tokens in the context.
Technical
Query · Key · Value
Query
“What information am I looking for?”
Comes from the token currently being computed, “it” in our example.
“What kind of information do I contain?”
Comes from every token in the sequence, including the query token itself.
Value
“What information can I provide?”
Also comes from every token; this is what actually gets passed on, weighted by the score.
This is an intuition, not a literal mathematical definition. Q, K, and V are simply three learned linear projections of the same input.
Step by step
Worked example: query = “it”, key = “robot”
The query vector for "it" is compared against the key vector for "robot" with a dot product, then scaled down by √dₖ to keep the numbers in a stable range.
An attention map is one observable piece of the computation, not a transcript of the model's reasoning. Treating it as a complete explanation of why a model produced a particular output overstates what it shows.
From one token to the whole sentence
The steps above followed a single pair, "it" and "robot". The model computes every pair at once with matrices. Stack the token vectors as the rows of (shape ) and multiply by three learned matrices:
Row of is the query of token , row of is the key of token , and row of is its value. and have columns and has (usually the same number). Where these three matrices come from is the next chapter's subject. For now, read them as three different views of the same tokens.
Scores
Comparing every query with every key is one matrix product, . Its entry in row , column is the dot product of query with key : multiply matching coordinates and add. It is large when the two vectors point in similar directions. Dividing by gives the scores:
This matrix is where the inspector's raw score comes from. The division deserves an explanation. If the entries of and are independent with mean 0 and variance 1, their dot product has variance , so raw scores grow as the vectors get longer. Large scores push softmax toward giving almost all the weight to one token, where its gradients become tiny and training slows down. Dividing by brings the variance back to 1 whatever the head size.
Weights
Softmax is applied to each row separately, so row becomes a set of positive numbers that sum to 1. Read it as a budget: token has one unit of attention to spend, and every share given to one token is taken from the others. The exponential makes the best matches stand out without setting the rest exactly to zero, so gradients can reach every entry during training.
Mixing the values
The output for token is a weighted average of all the value vectors, with row of as the weights. In the worked example, is mostly the value vector of "robot", plus smaller shares of the others. Keys decide who gets heard and values decide what they say. Because they are separate projections, a token can be easy to find for one reason and hand over something else entirely. Chaining the three lines gives the formula at the top of this page. The result has one row per token, the same as the input, and it is normally added to the token's own vector (chapter 10) rather than replacing it.
Shapes through one attention head
X (n×d) ─ ×W^Q → Q (n×d_k) ─┐
─ ×W^K → K (n×d_k) ─┴→ S = QKᵀ/√d_k (n×n scores)
│ softmax, row by row
▼
A (n×n weights)
│
─ ×W^V → V (n×d_v) ─────┴→ Z = A·V (n×d_v)The mask: what a text generator may not look at
The matrix in this chapter lets every token see every other token, including later ones. That is how encoder models such as BERT work: they read a whole text at once. The models this lab is about, GPT-style generators, predict the next token, and they can't be allowed to peek at it. So they add a mask before the softmax:
Since , a token gets exactly zero weight on anything after it, and each row renormalizes over what it can see. In a causal model the "it" row would spend its budget on the first seven tokens only. The mask has two practical payoffs. During training, every position can learn to predict its next token in a single pass (chapter 16). During generation, the keys and values of earlier tokens never change as new tokens arrive, so they can be cached (chapter 15). The original Transformer used both styles: bidirectional attention in its encoder and masked attention in its decoder.
One head, one layer
The heat map shows the weights , but what moves between tokens is the value vectors they select, and the layers after attention transform those further. Every head in every layer has its own matrix, so what you see here is one head in one layer (multi-head attention comes two chapters later).
One more property: the score matrix has entries, so compute and memory for attention grow with the square of the sequence length. That is a big reason long contexts are expensive, and why variants such as sliding-window attention exist, along with faster exact implementations such as FlashAttention. Next, the three matrices , and : what they are, and why there are three.