Feed Forward
After tokens exchange information, each one is transformed independently.
Attention and the feed-forward network do two very different jobs. Attention mixes information across tokens. It's the only place in a transformer block where tokens exchange information with each other. The feed-forward network transforms each token's representation entirely on its own, with no knowledge of any other token's value at that layer.
ReLU zeroed 3 of 8 hidden units for "robot" this pass: any pre-activation value below zero becomes exactly zero. Weights are fixed, illustrative values, not a trained model's.
The vector is expanded to a wider hidden dimension, passed through a nonlinear activation (ReLU above; some models use variants like GELU), then projected back down to the original width. The nonlinearity is what matters most: without it, stacking linear layers would collapse into one big linear layer, no matter how many you stacked.
What happens to each vector
After attention, each token's vector holds what it gathered from its context. The feed-forward network (FFN, also called an MLP) now works on that vector, one token at a time. It is two learned linear maps with a nonlinear function between them. Here is one token's vector of width , the first matrix has shape , the second has shape , and is an activation function applied to each number separately. In the original Transformer, GPT-2 and GPT-3 the hidden width is : 3,072 for GPT-2 small and 49,152 for GPT-3. The demo above uses 4, 8, 4 so that every number fits on screen.
Expand, switch, compress
Each of the hidden units is a small detector. Its column in is a pattern to look for, the dot product measures how well the incoming vector matches it, and the activation decides how strongly the unit responds. ReLU, the activation in the demo, is : a poor match is silenced completely. The second matrix then writes the result. Every unit that fired adds its own row of to the output, scaled by how active it was:
Read this way, the layer is a bank of pattern detectors that each contribute a stored vector when they fire. Researchers studying trained models have described it as a kind of key-value memory, and some of a model's factual knowledge (which city goes with which country, say) appears to be stored in these layers. Treat that as a useful working picture from interpretability research, not as a settled account of where any particular fact lives.
Why the nonlinearity is essential
Without the two matrices collapse into one: . Stacking ninety-six such layers would still be a single linear map, and a linear map can't express something like "respond only if A and B are both present but C is not". The activation removes that collapse, which is what lets extra depth add power. ReLU is the simplest choice. GPT-2 and BERT use GELU, a smooth version of the same idea. Many recent models, including the LLaMA family, use a gated variant called SwiGLU, which has three matrices instead of two:
The element-wise product lets one branch act as a gate on the other. To keep the parameter count comparable, these models usually shrink the hidden width to about . Many of them also leave out the bias terms and .
The same function at every position
The same and are applied to every token in the sequence, and no token can see another inside the FFN: each row of the matrix is processed on its own. The division of labor is clean. Attention moves information between tokens, and the FFN transforms what each token now holds. A block needs both. Gathering without processing would only average vectors together, and processing without gathering would leave every token blind to its context.
With the FFN has weights, twice the in the four attention matrices, so about two thirds of a standard block's parameters sit here. Some recent models replace the single FFN with many smaller "expert" FFNs and send each token through only a few of them, which raises total capacity without raising the work done per token.
The FFN's output is not passed on as it is. It is added back onto the vector that entered the sub-layer, and normalization keeps the scale steady. The next chapter assembles attention, the FFN, those additions and the normalization into one transformer block.