06
Attention stage

Attention

Each token gathers information from the other tokens around it, weighting each by how relevant it is.

By now each token has a vector that says which token it is and where it sits, and those vectors have never interacted. Yet meaning depends on other words: "bank" means something different next to "river" than next to "loan", and the "it" in the sentence below points back to someone. Attention is the operation that lets each token's vector be updated with information from other tokens. It is the only place in a transformer block where information moves between tokens.

Example sentence: “The robot carried the patient because it was tired.” It’s deliberately ambiguous: who does "it" refer to?

Therobotcarriedthepatientbecauseitwastired.Therobotcarriedthepatientbecauseitwastired.The → The: 0.54The → robot: 0.20The → carried: 0.04The → the: 0.03The → patient: 0.03The → because: 0.03The → it: 0.03The → was: 0.03The → tired: 0.03The → .: 0.03robot → The: 0.10robot → robot: 0.57robot → carried: 0.08robot → the: 0.03robot → patient: 0.04robot → because: 0.03robot → it: 0.05robot → was: 0.03robot → tired: 0.03robot → .: 0.03carried → The: 0.03carried → robot: 0.29carried → carried: 0.15carried → the: 0.03carried → patient: 0.38carried → because: 0.02carried → it: 0.03carried → was: 0.03carried → tired: 0.03carried → .: 0.02the → The: 0.04the → robot: 0.04the → carried: 0.04the → the: 0.17the → patient: 0.56the → because: 0.03the → it: 0.04the → was: 0.03the → tired: 0.03the → .: 0.03patient → The: 0.03patient → robot: 0.04patient → carried: 0.15patient → the: 0.07patient → patient: 0.51patient → because: 0.03patient → it: 0.04patient → was: 0.03patient → tired: 0.06patient → .: 0.03because → The: 0.04because → robot: 0.05because → carried: 0.07because → the: 0.04because → patient: 0.07because → because: 0.37because → it: 0.11because → was: 0.05because → tired: 0.18because → .: 0.03it → The: 0.04it → robot: 0.63it → carried: 0.04it → the: 0.04it → patient: 0.06it → because: 0.04it → it: 0.06it → was: 0.04it → tired: 0.04it → .: 0.03was → The: 0.02was → robot: 0.05was → carried: 0.03was → the: 0.02was → patient: 0.03was → because: 0.03was → it: 0.40was → was: 0.10was → tired: 0.28was → .: 0.03tired → The: 0.02tired → robot: 0.11tired → carried: 0.02tired → the: 0.02tired → patient: 0.05tired → because: 0.03tired → it: 0.23tired → was: 0.16tired → tired: 0.34tired → .: 0.02. → The: 0.03. → robot: 0.10. → carried: 0.04. → the: 0.03. → patient: 0.06. → because: 0.03. → it: 0.07. → was: 0.06. → tired: 0.25. → .: 0.33

Token inspector

Query (row)
it
Key (column)
robot
Raw score (QKᵀ/√d)
1.90
Attention weight
0.629

Click any cell. Row it's weights always sum to 1 across every column. That's exactly what a row of softmax output means.

These weights are hand-designed for this example, not the output of a trained model. Their only purpose is to make the shape of the computation easy to see. A real model learns attention patterns like these through training, never by hand.

Notice that “it” attends most strongly back to “robot,” not to “patient.” The sentence itself leaves that relationship ambiguous on the surface, yet attention can still surface a pattern like this, even though nothing in the sentence states it explicitly.

Intuitive

Attention is a mechanism that lets a token's representation take into account information from other tokens in the context.

Technical

Attention(Q,K,V)=softmax ⁣(QKTdk)V\text{Attention}(Q,K,V) = \text{softmax}\!\left(\frac{QK^T}{\sqrt{d_k}}\right)V

Query · Key · Value

Query

“What information am I looking for?”

Comes from the token currently being computed, “it” in our example.

Key

“What kind of information do I contain?”

Comes from every token in the sequence, including the query token itself.

Value

“What information can I provide?”

Also comes from every token; this is what actually gets passed on, weighted by the score.

This is an intuition, not a literal mathematical definition. Q, K, and V are simply three learned linear projections of the same input.

Step by step

Step by step

Worked example: query = “it”, key = “robot”

eit, robot=qit⋅krobotdke_{\text{it},\,\text{robot}} = \dfrac{q_{\text{it}} \cdot k_{\text{robot}}}{\sqrt{d_k}}

The query vector for "it" is compared against the key vector for "robot" with a dot product, then scaled down by √dₖ to keep the numbers in a stable range.

1.90raw score

An attention map is one observable piece of the computation, not a transcript of the model's reasoning. Treating it as a complete explanation of why a model produced a particular output overstates what it shows.

From one token to the whole sentence

The steps above followed a single pair, "it" and "robot". The model computes every pair at once with matrices. Stack the nn token vectors as the rows of XX (shape n×dn \times d) and multiply by three learned matrices:

Q=XWQ,K=XWK,V=XWVQ = XW^Q, \qquad K = XW^K, \qquad V = XW^V

Row ii of QQ is the query of token ii, row jj of KK is the key of token jj, and row jj of VV is its value. QQ and KK have dkd_k columns and VV has dvd_v (usually the same number). Where these three matrices come from is the next chapter's subject. For now, read them as three different views of the same tokens.

Scores

Comparing every query with every key is one matrix product, QKTQK^T. Its entry in row ii, column jj is the dot product of query ii with key jj: multiply matching coordinates and add. It is large when the two vectors point in similar directions. Dividing by dk\sqrt{d_k} gives the scores:

S=QKTdk,Sij=qi⋅kjdkS = \dfrac{QK^T}{\sqrt{d_k}}, \qquad S_{ij} = \dfrac{q_i \cdot k_j}{\sqrt{d_k}}

This n×nn \times n matrix is where the inspector's raw score comes from. The division deserves an explanation. If the entries of qq and kk are independent with mean 0 and variance 1, their dot product has variance dkd_k, so raw scores grow as the vectors get longer. Large scores push softmax toward giving almost all the weight to one token, where its gradients become tiny and training slows down. Dividing by dk\sqrt{d_k} brings the variance back to 1 whatever the head size.

Weights

A=softmax(S),Aij=eSij∑j′eSij′A = \text{softmax}(S), \qquad A_{ij} = \dfrac{e^{S_{ij}}}{\sum_{j'} e^{S_{ij'}}}

Softmax is applied to each row separately, so row ii becomes a set of positive numbers that sum to 1. Read it as a budget: token ii has one unit of attention to spend, and every share given to one token is taken from the others. The exponential makes the best matches stand out without setting the rest exactly to zero, so gradients can reach every entry during training.

Mixing the values

Z=AV,zi=∑jAij vjZ = AV, \qquad z_i = \sum_{j} A_{ij}\, v_j

The output for token ii is a weighted average of all the value vectors, with row ii of AA as the weights. In the worked example, zitz_{\text{it}} is mostly the value vector of "robot", plus smaller shares of the others. Keys decide who gets heard and values decide what they say. Because they are separate projections, a token can be easy to find for one reason and hand over something else entirely. Chaining the three lines gives the formula at the top of this page. The result has one row per token, the same as the input, and it is normally added to the token's own vector (chapter 10) rather than replacing it.

Shapes through one attention head

X (n×d) ─ ×W^Q → Q (n×d_k) ─┐
        ─ ×W^K → K (n×d_k) ─┴→ S = QKᵀ/√d_k   (n×n scores)
                                 │ softmax, row by row
                                 ▼
                             A  (n×n weights)
                                 │
        ─ ×W^V → V (n×d_v) ─────┴→ Z = A·V      (n×d_v)

The mask: what a text generator may not look at

The matrix in this chapter lets every token see every other token, including later ones. That is how encoder models such as BERT work: they read a whole text at once. The models this lab is about, GPT-style generators, predict the next token, and they can't be allowed to peek at it. So they add a mask before the softmax:

A=softmax ⁣(QKTdk+M),Mij={0j≤i−∞j>iA = \text{softmax}\!\left(\dfrac{QK^T}{\sqrt{d_k}} + M\right), \qquad M_{ij} = \begin{cases} 0 & j \le i \\ -\infty & j > i \end{cases}

Since e−∞=0e^{-\infty} = 0, a token gets exactly zero weight on anything after it, and each row renormalizes over what it can see. In a causal model the "it" row would spend its budget on the first seven tokens only. The mask has two practical payoffs. During training, every position can learn to predict its next token in a single pass (chapter 16). During generation, the keys and values of earlier tokens never change as new tokens arrive, so they can be cached (chapter 15). The original Transformer used both styles: bidirectional attention in its encoder and masked attention in its decoder.

One head, one layer

The heat map shows the weights AA, but what moves between tokens is the value vectors they select, and the layers after attention transform those further. Every head in every layer has its own matrix, so what you see here is one head in one layer (multi-head attention comes two chapters later).

One more property: the score matrix has n×nn \times n entries, so compute and memory for attention grow with the square of the sequence length. That is a big reason long contexts are expensive, and why variants such as sliding-window attention exist, along with faster exact implementations such as FlashAttention. Next, the three matrices WQW^Q, WKW^K and WVW^V: what they are, and why there are three.

Key Takeaway

Attention lets each token gather information from the other tokens it can see, weighted by how relevant each one is. That weighting comes from queries, keys, and values, then gets normalized with softmax so it forms a proper distribution.