04
Embedding stage

Embeddings

Each token becomes a list of numbers that places it in a space of meaning.

Token IDs are labels. Adding two of them, or asking which is closer to a third, produces numbers with no meaning. To calculate with tokens, the model needs a representation in which distance and direction can carry information, and one that can be adjusted so training can improve it. A vector of real numbers satisfies both.

Choose a word

Human view

cat

animalsmallfurrylivingdomesticated

Machine view

[-0.53, 0.26, -0.06, 0.86, 0.51, 0.24, 0.55, 0.66, …]

A 2D projection of embedding space

catdogliontigercartruckbicyclepatientdoctornurserobotcomputermachine

Real embeddings live in high-dimensional spaces (often hundreds or thousands of dimensions). This 2D layout is a hand-placed projection for human intuition, not a mathematical reduction of a real model's weights. Notice, though, how words with related meaning still end up near each other.

What is happening?

Once text is tokenized, each token id is looked up in a table and replaced with a vector, a fixed-length list of numbers. That table is learned during training: tokens that tend to appear in similar contexts end up with similar vectors, which is why "cat" and "dog" land near each other above while "car" sits somewhere else entirely.

This is a genuinely different kind of representation from a dictionary definition. Nothing in the vector is labeled "animal" or "vehicle." Those clusters emerge from the numbers' relative positions; nothing writes them into any single coordinate.

A table with one row per token

The embedding matrix EE has one row for each token in the vocabulary and dd columns, where dd is the model's width: 768 in GPT-2 small, 4,096 in a 7-billion-parameter LLaMA-family model, 12,288 in GPT-3. Embedding a token means reading off its row:

xi=E[ti],E∈R∣V∣×dx_i = E[t_i], \qquad E \in \mathbb{R}^{|V| \times d}

Here tit_i is the ID of the ii-th token and xix_i is a vector of dd numbers. The same thing can be written as multiplying a one-hot vector (zeros everywhere except a 1 at position tit_i) by EE, which is why the embedding counts as an ordinary layer of the network. Its size is easy to work out: 50,257×76850{,}257 \times 768 is about 38.6 million parameters for GPT-2 small.

Nobody fills this table in. It starts out random and is adjusted by gradient descent (chapter 16) along with every other parameter. The clusters in the projection above aren't designed either: tokens that can be swapped into the same contexts are pushed toward similar rows, because that helps predict what comes next.

Reading an embedding

The demo shows two views of the same word. The human view is a handful of tags we wrote, and it is there only to help you read the picture. The machine view is what the model actually has: a list of numbers (8 in this demo, hundreds or thousands in a real model) with no labels on any of them. A single coordinate rarely means anything by itself. Meaning is spread across many coordinates at once, so it lives in directions and distances rather than in individual columns.

The usual way to compare two embeddings is cosine similarity:

cos⁡(u,v)=u⋅v∥u∥ ∥v∥\cos(u, v) = \dfrac{u \cdot v}{\lVert u \rVert \, \lVert v \rVert}

It is 1 when the vectors point the same way, 0 when they are unrelated (perpendicular), and −1-1 when they point in opposite directions. The 2D picture is a hand-placed shadow of that geometry, so read it as a sketch: it shows that related words sit near each other, not any exact distances.

One vector per token, before any context

The lookup depends only on the token's ID. The word "bank" gets exactly the same vector in "river bank" and in "bank account", and two occurrences of "the" in one sentence get identical vectors. The embedding also says nothing about where in the sentence the token sits. At this point the model knows which token each one is, and nothing more.

That limitation explains the chapters ahead. Position has to be added (next chapter), and the words around each token have to be brought in (attention, from chapter 6). After the first few blocks, the vector at "bank" is no longer the raw table entry: it has been rewritten using its neighbours, which is what makes it a contextual representation.

Stacking the nn token vectors gives a matrix XX of shape n×dn \times d, one row per token. That matrix is what flows through the rest of the network. The same table EE often returns at the very end, used in reverse to score every token in the vocabulary (chapter 12).

Key Takeaway

An embedding is a learned vector for a token: not a definition, but a position in a space where distance and direction reflect something about usage and meaning.