Transformer Stack
Dozens of transformer blocks are stacked, each refining the representation further.
A single transformer block does one round of "mix information across tokens, then transform each token." A real model stacks many of these blocks (a few dozen is typical for a mid-sized model), feeding each block's output directly into the next block's input.
Block 6 of 12
Middle of the stack. Representations here have had room to combine information from further across the sequence, often blending local detail with broader context, such as which earlier phrase a pronoun points back to or how one clause relates to the rest of the sentence.
Every block shares the same internal structure. The Transformer Block chapter covers what happens inside each one.
Descriptions of "what early layers do" versus "what late layers do" are broad tendencies, observed in some models under some interpretability methods, not an architectural guarantee. Nothing in the transformer's design assigns a fixed job to any particular layer; whatever specialization emerges is a product of training.
Same shape in, same shape out
Every block maps an matrix to another matrix, so blocks can be chained. With the embeddings plus position signal and the number of blocks:
Each block has its own weights. Block 7 is not block 3 run again; it has learned a different job, which is why a model with more layers is more capable and more expensive, rather than merely the same thing repeated. The lab's diagram shows , the depth of GPT-2 small. Real sizes vary widely:
A few real configurations
model layers width d heads parameters GPT-2 small 12 768 12 124 million LLaMA 2 7B 32 4,096 32 6.7 billion GPT-3 96 12,288 96 175 billion
The residual stream
Because each sub-layer adds to its input instead of replacing it, there is a simpler way to see the whole stack. Each token has a running vector, the residual stream, that starts as its embedding and is written to by every sub-layer in turn. In the pre-norm layout a block's result is its input plus what attention added plus what the FFN added, so after blocks:
Here is what attention in block added to this token and is what its FFN added. Each sub-layer reads the stream, computes something, and writes a correction back. Nothing is overwritten by design; a sub-layer that wants to erase something has to write its opposite. That is why information can travel from the first layer to the last without anyone passing it along explicitly, and why later layers can build on what earlier ones left behind.
What depth buys
One attention layer lets a token gather once. A second layer can gather from tokens that have already gathered, so it can combine information two steps away: a pronoun finds its noun in one layer, and a later layer can use what that noun had itself collected earlier. Depth is how a model goes from single relationships to composed ones.
Where the parameters are
Each block holds about weights (previous chapter) and the embedding table holds . For GPT-3 that is billion plus about 0.6 billion for embeddings, close to the advertised 175 billion. The formula assumes a standard layout (four attention matrices and a feed-forward layer of width ), so models with grouped-query attention or different feed-forward shapes deviate from it.
From the last block to a prediction
The output of block is an matrix of hidden representations (often called hidden states): one vector per token, enriched by every layer. In a pre-norm model it passes through one more LayerNorm first. To predict the next token only the last position's vector is needed, because that token has attended to the entire context. The next chapter turns that vector into a score for every token in the vocabulary.