09
Attention stage

Feed Forward

After tokens exchange information, each one is transformed independently.

Attention and the feed-forward network do two very different jobs. Attention mixes information across tokens. It's the only place in a transformer block where tokens exchange information with each other. The feed-forward network transforms each token's representation entirely on its own, with no knowledge of any other token's value at that layer.

Input
× W₁ + b₁
Pre-activation
ReLU
Hidden
× W₂ + b₂
Output

ReLU zeroed 3 of 8 hidden units for "robot" this pass: any pre-activation value below zero becomes exactly zero. Weights are fixed, illustrative values, not a trained model's.

FFN(x)=σ(xW1+b1) W2+b2\text{FFN}(x) = \sigma(x W_1 + b_1)\,W_2 + b_2

The vector is expanded to a wider hidden dimension, passed through a nonlinear activation (ReLU above; some models use variants like GELU), then projected back down to the original width. The nonlinearity is what matters most: without it, stacking linear layers would collapse into one big linear layer, no matter how many you stacked.

What happens to each vector

After attention, each token's vector holds what it gathered from its context. The feed-forward network (FFN, also called an MLP) now works on that vector, one token at a time. It is two learned linear maps with a nonlinear function between them. Here xx is one token's vector of width dd, the first matrix has shape d×dffd \times d_{ff}, the second has shape dff×dd_{ff} \times d, and σ\sigma is an activation function applied to each number separately. In the original Transformer, GPT-2 and GPT-3 the hidden width is dff=4dd_{ff} = 4d: 3,072 for GPT-2 small and 49,152 for GPT-3. The demo above uses 4, 8, 4 so that every number fits on screen.

Expand, switch, compress

Each of the dffd_{ff} hidden units is a small detector. Its column wkw_k in W1W_1 is a pattern to look for, the dot product measures how well the incoming vector matches it, and the activation decides how strongly the unit responds. ReLU, the activation in the demo, is max⁡(0,z)\max(0, z): a poor match is silenced completely. The second matrix then writes the result. Every unit that fired adds its own row uku_k of W2W_2 to the output, scaled by how active it was:

ak=σ(x⋅wk+b1,k),FFN(x)=∑k=1dffak uk+b2a_k = \sigma\big(x \cdot w_k + b_{1,k}\big), \qquad \text{FFN}(x) = \sum_{k=1}^{d_{ff}} a_k\, u_k + b_2

Read this way, the layer is a bank of pattern detectors that each contribute a stored vector when they fire. Researchers studying trained models have described it as a kind of key-value memory, and some of a model's factual knowledge (which city goes with which country, say) appears to be stored in these layers. Treat that as a useful working picture from interpretability research, not as a settled account of where any particular fact lives.

Why the nonlinearity is essential

Without σ\sigma the two matrices collapse into one: (xW1)W2=x(W1W2)(xW_1)W_2 = x(W_1W_2). Stacking ninety-six such layers would still be a single linear map, and a linear map can't express something like "respond only if A and B are both present but C is not". The activation removes that collapse, which is what lets extra depth add power. ReLU is the simplest choice. GPT-2 and BERT use GELU, a smooth version of the same idea. Many recent models, including the LLaMA family, use a gated variant called SwiGLU, which has three matrices instead of two:

FFNSwiGLU(x)=(SiLU(xWg)⊙xWu) Wd,SiLU(z)=z1+e−z\text{FFN}_{\text{SwiGLU}}(x) = \big(\text{SiLU}(xW_g) \odot xW_u\big)\,W_d, \qquad \text{SiLU}(z) = \dfrac{z}{1 + e^{-z}}

The element-wise product ⊙\odot lets one branch act as a gate on the other. To keep the parameter count comparable, these models usually shrink the hidden width to about 83d\tfrac{8}{3}d. Many of them also leave out the bias terms b1b_1 and b2b_2.

The same function at every position

The same W1W_1 and W2W_2 are applied to every token in the sequence, and no token can see another inside the FFN: each row of the n×dn \times d matrix is processed on its own. The division of labor is clean. Attention moves information between tokens, and the FFN transforms what each token now holds. A block needs both. Gathering without processing would only average vectors together, and processing without gathering would leave every token blind to its context.

With dff=4dd_{ff} = 4d the FFN has 2⋅d⋅4d=8d22 \cdot d \cdot 4d = 8d^2 weights, twice the 4d24d^2 in the four attention matrices, so about two thirds of a standard block's parameters sit here. Some recent models replace the single FFN with many smaller "expert" FFNs and send each token through only a few of them, which raises total capacity without raising the work done per token.

The FFN's output is not passed on as it is. It is added back onto the vector that entered the sub-layer, and normalization keeps the scale steady. The next chapter assembles attention, the FFN, those additions and the normalization into one transformer block.

Key Takeaway

The feed-forward network transforms each token independently: expand, apply a nonlinearity, project back down. It's where per-token computation happens, complementing attention's cross-token mixing.