Position
Order matters, so the model adds a signal for where each token sits in the sequence.
Embeddings gave every token a vector, and each vector depends only on which token it is. Take the six vectors for this sentence and shuffle them, and nothing inside them would reveal the shuffle. Whether "dog bites man" or "man bites dog" was written is a question about order, and so far order isn't stored anywhere.
Why position needs its own signal
Attention, which the next chapter covers, treats a sentence more like a set than a sequence. On its own, it has no built-in sense of which token came first. "The robot follows the patient" and a shuffled version of the same words would look identical to it unless order is encoded somewhere. So before anything else happens, a signal for each token's position is combined with its embedding.
One common signal: sinusoidal encoding
One widely used approach adds a wave to each dimension of the embedding, with a different frequency per dimension, so that every position ends up with its own unique pattern across all of them.
Horizontal axis is token position (0 to 23). Each dimension pair oscillates at its own frequency: pair 0 repeats about every 6 positions, pair 4 about every 63, and pair 12 barely moves across these 24 positions. Together they give every position a unique combination of values.
This sinusoidal form is one specific, historically important choice, not the only one. Many current models instead learn the position signal directly, or bake relative position into attention itself. The wave-based version is used here because its structure is easy to see directly, as in the chart above.
Reading the formulas
With model width , the position signal for position is another vector of numbers, built from pairs. Pair number uses one angle, , and contributes its sine to dimension and its cosine to dimension . Only the frequency changes from pair to pair, and the signal is added to the embedding:
The first pair () has , so it completes a cycle about every 6 tokens. Each later pair is slower, and the slowest ones take tens of thousands of tokens to cycle once. Think of an odometer: the last digit spins fast and tells neighbouring positions apart, the leading digits move slowly and separate distant ones, and together the digits give every position its own reading. In the chart above, the third wave (pair 12 of a 32-wide model) has a wavelength of about 6,300 tokens, so across these 24 positions it barely leaves zero. That is by design: it only starts to matter over long distances.
Why sine and cosine come in pairs
The sine and cosine of one angle form a point on a circle, and moving positions forward turns that point by a fixed angle , wherever in the sequence you start:
So "k steps later" is the same simple transformation at every position, something a network can learn to detect with ordinary matrix multiplications. The original Transformer paper chose this design partly for that reason. It is also the idea that rotary embeddings, below, take further.
Why it is added, not attached
The two vectors are summed, so the width stays and nothing else in the network has to change. That sounds like it should blur a token's identity with its position. In a space with hundreds or thousands of dimensions the two signals mostly occupy different directions, and the network learns to read both from the sum.
Other ways to encode position
- Learned absolute positions. GPT-2 keeps a second table with one trainable vector for every position up to its maximum length and adds it just like the sinusoidal signal. It works well within the trained length, but it has no vectors for positions it never saw.
- Rotary embeddings (RoPE). Instead of adding anything to the embedding, RoPE rotates each query and key vector by an angle proportional to its position. The dot product of a query and a key then depends on the content of the two tokens and on how far apart they are, not on where they sit in absolute terms. Many recent open-weight models, including the LLaMA family, use it.
- ALiBi. Adds nothing to the embeddings. It subtracts a penalty from each attention score that grows with the distance between the two tokens.
Each token vector now carries what the token is and where it sits. That combined matrix is the input to the first transformer block, and inside it, attention is the step where the tokens start to influence each other.