05
Embedding stage

Position

Order matters, so the model adds a signal for where each token sits in the sequence.

Embeddings gave every token a vector, and each vector depends only on which token it is. Take the six vectors for this sentence and shuffle them, and nothing inside them would reveal the shuffle. Whether "dog bites man" or "man bites dog" was written is a question about order, and so far order isn't stored anywhere.

Theposition 0
robotposition 1
followsposition 2
theposition 3
patientposition 4
.position 5

Why position needs its own signal

Attention, which the next chapter covers, treats a sentence more like a set than a sequence. On its own, it has no built-in sense of which token came first. "The robot follows the patient" and a shuffled version of the same words would look identical to it unless order is encoded somewhere. So before anything else happens, a signal for each token's position is combined with its embedding.

Token embedding+Position signal=Transformer input

One common signal: sinusoidal encoding

One widely used approach adds a wave to each dimension of the embedding, with a different frequency per dimension, so that every position ends up with its own unique pattern across all of them.

dimension pair i = 0dimension pair i = 4dimension pair i = 12

Horizontal axis is token position (0 to 23). Each dimension pair oscillates at its own frequency: pair 0 repeats about every 6 positions, pair 4 about every 63, and pair 12 barely moves across these 24 positions. Together they give every position a unique combination of values.

Even dimensions
PE(pos, 2i)=sin⁡ ⁣(pos100002i/d)PE_{(pos,\,2i)} = \sin\!\left(\dfrac{pos}{10000^{2i/d}}\right)
Odd dimensions
PE(pos, 2i+1)=cos⁡ ⁣(pos100002i/d)PE_{(pos,\,2i+1)} = \cos\!\left(\dfrac{pos}{10000^{2i/d}}\right)

This sinusoidal form is one specific, historically important choice, not the only one. Many current models instead learn the position signal directly, or bake relative position into attention itself. The wave-based version is used here because its structure is easy to see directly, as in the chart above.

Reading the formulas

With model width dd, the position signal for position pospos is another vector of dd numbers, built from pairs. Pair number ii uses one angle, pos⋅ωipos \cdot \omega_i, and contributes its sine to dimension 2i2i and its cosine to dimension 2i+12i+1. Only the frequency ωi\omega_i changes from pair to pair, and the signal is added to the embedding:

ωi=110000 2i/d,zpos=xpos+PEpos\omega_i = \dfrac{1}{10000^{\,2i/d}}, \qquad z_{pos} = x_{pos} + PE_{pos}

The first pair (i=0i = 0) has ω0=1\omega_0 = 1, so it completes a cycle about every 6 tokens. Each later pair is slower, and the slowest ones take tens of thousands of tokens to cycle once. Think of an odometer: the last digit spins fast and tells neighbouring positions apart, the leading digits move slowly and separate distant ones, and together the digits give every position its own reading. In the chart above, the third wave (pair 12 of a 32-wide model) has a wavelength of about 6,300 tokens, so across these 24 positions it barely leaves zero. That is by design: it only starts to matter over long distances.

Why sine and cosine come in pairs

The sine and cosine of one angle form a point on a circle, and moving kk positions forward turns that point by a fixed angle kωik\omega_i, wherever in the sequence you start:

(sin⁡((pos+k) ωi)cos⁡((pos+k) ωi))=(cos⁡kωisin⁡kωi−sin⁡kωicos⁡kωi)(sin⁡(pos ωi)cos⁡(pos ωi))\begin{pmatrix} \sin\big((pos+k)\,\omega_i\big) \\ \cos\big((pos+k)\,\omega_i\big) \end{pmatrix} = \begin{pmatrix} \cos k\omega_i & \sin k\omega_i \\ -\sin k\omega_i & \cos k\omega_i \end{pmatrix} \begin{pmatrix} \sin\big(pos\,\omega_i\big) \\ \cos\big(pos\,\omega_i\big) \end{pmatrix}

So "k steps later" is the same simple transformation at every position, something a network can learn to detect with ordinary matrix multiplications. The original Transformer paper chose this design partly for that reason. It is also the idea that rotary embeddings, below, take further.

Why it is added, not attached

The two vectors are summed, so the width stays dd and nothing else in the network has to change. That sounds like it should blur a token's identity with its position. In a space with hundreds or thousands of dimensions the two signals mostly occupy different directions, and the network learns to read both from the sum.

Other ways to encode position

  • Learned absolute positions. GPT-2 keeps a second table with one trainable vector for every position up to its maximum length and adds it just like the sinusoidal signal. It works well within the trained length, but it has no vectors for positions it never saw.
  • Rotary embeddings (RoPE). Instead of adding anything to the embedding, RoPE rotates each query and key vector by an angle proportional to its position. The dot product of a query and a key then depends on the content of the two tokens and on how far apart they are, not on where they sit in absolute terms. Many recent open-weight models, including the LLaMA family, use it.
  • ALiBi. Adds nothing to the embeddings. It subtracts a penalty from each attention score that grows with the distance between the two tokens.

Each token vector now carries what the token is and where it sits. That combined matrix is the input to the first transformer block, and inside it, attention is the step where the tokens start to influence each other.

Key Takeaway

Attention has no built-in notion of order, so a position signal is added to every token's embedding before the transformer sees it. Sinusoidal encoding is one common way to build that signal.