Tokenization
Text is cut into tokens before anything mathematical can happen to it.
The previous chapter ended with characters: integers with no meaning attached. A model needs a better unit than that. It has to be small enough that any text can be built from it, and large enough that one unit carries something worth learning about. Tokens are that unit.
Tokens
Token IDs
| Position | Token | Pseudo ID |
|---|---|---|
| 0 | Trans | 41493 |
| 1 | ##form | 20529 |
| 2 | ##ers | 16889 |
| 3 | are | 21259 |
| 4 | power | 46099 |
| 5 | ##ful | 22514 |
| 6 | . | 26776 |
Educational tokenizer: simplified for visualization. It is not the tokenizer used by GPT, Claude, Gemini, or any other production model.
What is happening?
A tokenizer breaks text into pieces from a fixed vocabulary, then looks each piece up to get a token id, the integer the model actually operates on. A few things about that are easy to miss:
- A token is not always a word. Try a longer or less common word above; it often splits into several pieces (shown lighter, prefixed with ##).
- Punctuation gets its own tokens: a period is a token just as much as a word is.
- Tokenization is a design choice, not a law of nature. Different tokenizers split the same sentence differently, and produce different ids for the same word.
Every model has to make this trade-off somewhere: fewer, larger tokens make sequences shorter but the vocabulary bigger; more, smaller tokens do the opposite. There's no single correct split, only ones that work better or worse for a given model and language.
Three ways to cut text
Cutting text into single characters keeps the vocabulary tiny (a few hundred symbols) but makes sequences long, and the model has to spend its layers reassembling letters into words. Cutting into whole words gives short sequences but an enormous vocabulary, and anything outside it (a new name, a typo, "robotics" when only "robot" was ever seen) has no entry at all. Subword tokens sit in between: frequent words get a token of their own, and rarer words are assembled from smaller pieces that are already in the vocabulary. Nearly every current LLM works this way.
Where the pieces come from
The most common recipe is byte-pair encoding (BPE). It runs once, on a large body of text, before the model is trained. Start with single characters, count which adjacent pair occurs most often, merge that pair into a new symbol, and repeat until the vocabulary reaches the size you want. Frequent patterns end up as single tokens without anyone listing them by hand.
A tiny BPE run
corpus: low ×5 lower ×2 newest ×6 widest ×3 merge 1: e + s → es (occurs 9 times) merge 2: es + t → est (occurs 9 times) merge 3: l + o → lo (occurs 7 times) merge 4: lo + w → low (occurs 7 times) unseen word "lowest" → low | est
The tokenizer is then frozen. It is a separate program, not part of the neural network, and it never changes while the model trains. Real tokenizers usually work on bytes rather than characters (byte-level BPE), so any string in any script, emoji included, can be encoded without an "unknown" symbol.
What the demo simplifies
The tokenizer above is a stand-in. It cuts long words into fixed-size chunks so that subword splitting is easy to see, and it marks continuation pieces with ##, the convention of WordPiece tokenizers such as BERT's. GPT-style byte-level tokenizers don't use ##. They attach the leading space to the start of the next word instead, so " robot" (with a space) and "robot" are two different tokens.
From pieces to IDs
Every piece in the vocabulary has an integer ID, which is its row number in the tokenizer's table. The model's input is therefore not text but a list of IDs:
Here is the number of tokens and is the vocabulary size: about 50,000 for GPT-2, and over 100,000 for several recent models. An ID is an address, not a measurement. Token 4,113 is not "bigger" than token 4,112, and neighbouring IDs are unrelated. The pseudo IDs in the demo are hashes of each piece, so a piece always gets the same number; in a real tokenizer they follow the order in which pieces were added while the vocabulary was built. Special tokens, which mark the start of a text, its end, or the boundary between chat messages, have IDs too.
What the choice costs
The model only ever sees IDs, so whatever the tokenizer does poorly, the model inherits. Text that is rare in the tokenizer's training data (many non-English languages, unusual code formatting, long numbers) tends to be split into more, smaller pieces, so the same content takes more tokens. Indonesian, for example, often needs more tokens per word than English with a tokenizer built mostly from English text. That matters for cost, because context windows and prices are counted in tokens rather than words, and it can matter for quality. It is also one reason letter-level tasks, like counting the r's in a word, can trip a model up: it sees one token, not the letters inside it.
A list of integers still can't be added or compared in any meaningful way. The next chapter fixes that by giving every ID its own vector of learned numbers.