Transformer Block
Attention and feed-forward layers combine with residual connections into one repeatable unit.
Multi-head attention and the feed-forward network, covered in the last two chapters, are combined here into one repeatable unit, alongside two ingredients that rarely get attention on their own but matter enormously in practice: layer normalization and residual connections.
Multi-Head Attention
Tokens exchange information with each other here: each one gathers a weighted mixture of the other tokens' value vectors, using the query, key, and value mechanism from the Attention chapter.
The bracketed "+" paths are the residual connections: they carry a copy of the block's input forward, unchanged, to be added back in after attention and after the feed-forward network. Click any box to read what it does.
Why the residual connections matter
Without them, a token's representation would have to pass entirely through attention, then entirely through the feed-forward network, at every single block. With dozens of blocks stacked on top of each other, that makes training very unstable. A residual connection lets each sub-layer learn a small adjustment to add to the input, rather than having to reconstruct the whole representation from scratch, so most of the original signal survives intact all the way through the stack.
LayerNorm's job is narrower than it might seem: for each token on its own, it subtracts the average of the vector's numbers and divides by their spread, then applies a learned gain and offset. That keeps every sub-layer's input in a consistent range, which alone makes deep networks noticeably easier to train. It is not lossless: the vector's overall magnitude and average are removed, and the learned gain and offset only partly restore them. The pattern of values across its features survives. Placing LayerNorm before each sub-layer, as shown here, is usually called a pre-norm arrangement. It is the placement most modern large language models use, in place of the original 2017 Transformer paper's post-norm design, which normalized after each sub-layer instead.
Two sub-layers, wrapped the same way
A block has two sub-layers, attention and then the feed-forward network, and each one is wrapped in the same two helpers: a normalization before it and a residual connection around it. In the pre-norm arrangement drawn above (used by GPT-2 and most models since), a block takes the matrix of token vectors and computes:
, and all have shape . LN and the FFN act on each row separately, and attention is the only step that mixes rows. Because the output has the same shape as the input, the next block can take directly, which is what makes stacking possible.
Residual connections: learn the change, not the whole
Without the skip, a sub-layer has to produce its entire output from scratch. With it, the sub-layer only has to produce an adjustment. If the best adjustment is "nothing", the weights just need to make small and the token passes through intact. The derivative shows a second benefit. The identity term gives the training signal a direct path from the loss back to every earlier layer, so gradients don't have to survive a trip through dozens of multiplications. The idea comes from residual networks for images and the Transformer adopted it wholesale. Unrolled across a whole stack, each token's final vector is its starting embedding plus everything every sub-layer added, a view the next chapter calls the residual stream.
Layer normalization: keeping the scale steady
For one token, LayerNorm computes the mean and variance of the numbers in its vector, shifts and scales them to mean 0 and variance 1, then applies a learned gain and offset for each feature. The statistics come from that one vector alone, never from other tokens or other examples, so it behaves identically in training and generation. The small only prevents division by zero. The reason to bother is that every sub-layer adds something to the stream, so the vectors' scale can drift from block to block, and a sub-layer tuned for one scale misbehaves at another. Normalizing its input keeps each sub-layer in a range it was trained for. Many recent models use RMSNorm, a cheaper variant that skips the mean subtraction and divides by before applying the gain.
Pre-norm versus post-norm
The original 2017 Transformer normalized after each addition, , which puts LayerNorm on the main path so every gradient has to pass through it. Pre-norm keeps the main path clean and is reported to be easier to train at depth, at the cost of a stream whose scale tends to grow with depth. That is why pre-norm models apply one more LayerNorm after the last block, before producing scores. Both layouts exist in practice. This lab draws the pre-norm one.
What a block holds
The two sub-layers together hold about weights per block; the normalization adds only per LayerNorm. For GPT-2 small, where , that is about 7.1 million per block. Nothing in a block is specific to one position or one sentence: it is one fixed recipe, applied to every token.
One block lets every token gather from its context once and then process what it gathered. Handling syntax, coreference and world knowledge together takes many rounds of that, and supplying those rounds is the job of the stack.