02
Input stage

Text

A model never sees a sentence; it sees numbers standing in for pieces of it.

Everything in this lab starts with a sentence you could type into a chat box. Before any of the model's machinery can act on it, that sentence has to become something a computer can calculate with. This chapter is about what the computer actually holds at that starting point, and why it isn't yet in a form a language model can use.

What you see

Every chapter in this lab returns to one running example: a short, unambiguous sentence you can watch move through the whole pipeline.

“The robot follows the patient.”

This is what you see: a string of readable characters.

What the model receives

A language model doesn't receive a sentence, and it doesn't receive "meaning" in any sense a person would recognize. It receives a sequence of numbers, each one standing in for a small piece of text. Everything that follows in this lab (embeddings, attention, the whole transformer) operates on those numbers, not on the words themselves.

Text→Characters→Tokens

Text is already numbers, just not useful ones

A computer never stored letters as letters. Each character has an integer code in the Unicode standard ("T" is 84, "h" is 104, "e" is 101, a space is 32), and those integers are saved as bytes. So text is numeric from the very beginning. The trouble is that the codes say nothing about meaning. "a" is 97 and "b" is 98, yet those two letters are no more related than "a" and "z". Two words with nearly the same meaning have no numerical resemblance at all.

One sentence, three views

"The robot follows the patient."

characters  T, h, e, space, r, o, b, o, t, …      (30 in all)
Unicode     84, 104, 101, 32, 114, 111, 98, 111, 116, …
tokens      The | robot | follows | the | patient | .

Characters are also a poor unit for a model to reason over. This short sentence is already 30 characters long, yet it holds only a few ideas: a robot, a patient, an act of following. A model that worked one character at a time would need a very long sequence to express a modest sentence, and it would have to rediscover in every layer that "r-o-b-o-t" belongs together. So a step sits between raw characters and the model: characters are grouped into larger chunks called tokens.

What the model actually receives

A language model is a function with a very specific signature. It takes a sequence of tokens and returns, for every token in its vocabulary, the probability that it comes next. Writing t1,…,tnt_1,\dots,t_n for the input tokens and θ\theta for all the numbers the model has learned (its parameters), that signature is:

Pθ(tn+1∣t1,…,tn)P_\theta\left(t_{n+1} \mid t_1, \dots, t_n\right)

Read it as: given the tokens so far, how likely is each possible next token? Every later chapter is a closer look at one piece of computing that probability. Nothing in the expression mentions words, grammar or meaning. Whatever the model "knows" about robots and patients lives inside θ\theta, and it got there through training (chapter 16), not because anyone wrote rules.

Two example sentences

The lab uses two sentences. The robot follows the patient. is deliberately plain and travels through the early stages: text, tokens, embeddings and position. The robot carried the patient because it was tired. is deliberately ambiguous (who was tired?) and takes over in the attention chapters, because a pronoun that could point at two nouns is a good test for a mechanism that connects distant words. Some numbers in the visualizations are hand-made so a mechanism is easy to see. Each chapter says so where that applies, and says so too when a number comes from a live calculation.

The next chapter, Tokenization, is where that grouping of characters into tokens actually happens, and where you can try it on a sentence of your own.

Key Takeaway

A model's raw material is numbers standing in for pieces of text, not the text itself and not its meaning. Tokenization is the first step that makes this possible.