Glossary
Every term used across the lab, gathered in one place and linked back to where it's visualized.
- Token
A single unit of text a model operates on: often a word piece, sometimes a whole word or a punctuation mark.
See it visualized- Tokenizer
The algorithm that splits text into tokens and maps each one to an integer id.
See it visualized- Embedding
A learned vector representation of a token, positioned so that related tokens end up near each other.
See it visualized- Context
The full sequence of tokens a model can see when producing its next prediction.
- Attention
A mechanism that lets each token's representation incorporate information from other tokens, weighted by relevance.
See it visualized- Query
One of three projections used in attention. It represents what a token is 'looking for' in other tokens.
See it visualized- Key
One of three projections used in attention. It represents what a token 'offers' when compared against a query.
See it visualized- Value
One of three projections used in attention. It carries the information that actually gets passed on, weighted by attention score.
See it visualized- Logit
A raw, unnormalized score the model assigns to a candidate next token, before softmax.
See it visualized- Softmax
A function that turns a list of raw scores into a probability distribution that sums to 1.
See it visualized- Temperature
A parameter that reshapes a probability distribution before sampling: lower values sharpen it, higher values flatten it.
See it visualized- Transformer
The neural network architecture built from alternating attention and feed-forward layers that underlies most modern language models.
See it visualized- Residual connection
A shortcut that adds a layer's input back to its output, which helps very deep networks train.
See it visualized- LayerNorm
A normalization step in a transformer block: it re-centers and rescales each token's vector using that vector's own mean and spread, keeping values in a stable numerical range.
See it visualized- Feed-forward network
A small per-token neural network applied after attention, transforming each token's representation independently.
See it visualized- Parameter
A learned numerical weight inside the model. Modern models have anywhere from millions to hundreds of billions of them.
- Gradient
A measure of how much a small change in a parameter would change the loss, used to decide how to update that parameter during training.
See it visualized- Loss
A single number measuring how wrong the model's predictions were on a batch of examples during training.
See it visualized- Backpropagation
The algorithm that computes gradients for every parameter in the network from a single loss value.
- Fine-tuning
Further training a pretrained base model, typically to make it follow instructions or match a particular style.
See it visualized- RAG
Retrieval-augmented generation: retrieving relevant documents and adding them to context before generating an answer.
See it visualized- Vector database
A database optimized for finding the nearest embedding vectors to a query vector, used to power retrieval.
- Inference
Running a trained model to produce an output, as opposed to training it.
- Sampling
Choosing the next token from a probability distribution, not always the single most likely one.
See it visualized- Causal mask
A rule inside attention that stops each token from seeing the tokens after it, so a text generator can't peek at the answer. It is applied by adding −∞ to the forbidden scores before softmax, which turns their weights into exactly zero.
See it visualized- Positional encoding
A signal that tells the model where each token sits in the sequence, since attention alone has no sense of order. It can be added to the embeddings (sinusoidal or learned) or applied inside attention (rotary embeddings, ALiBi).
See it visualized- Multi-head attention
Several attention computations run in parallel, each with its own query, key and value projections in a smaller subspace. Their outputs are concatenated and mixed by one more learned matrix.
See it visualized- Dot product
Multiply two vectors coordinate by coordinate and add up the results. The number is large when the vectors point in similar directions, which is why it serves as the similarity score in attention.
See it visualized- Cosine similarity
A similarity measure that depends only on the angle between two vectors: 1 for the same direction, 0 for unrelated, −1 for opposite. Commonly used to compare embeddings, including in retrieval.
See it visualized- Vocabulary
The fixed set of tokens a model can read and write, each with an integer ID. Sizes range from tens of thousands to a few hundred thousand entries.
See it visualized- Hidden state
A token's vector at some point inside the network, between the embedding and the output scores. The one after the last block is what gets projected into logits.
See it visualized- Residual stream
The running vector each token carries through the stack. Every attention and feed-forward sub-layer reads it and adds its result back, so the final vector is the original embedding plus everything added along the way.
See it visualized- Unembedding
The final matrix that turns a token's last vector into one score per vocabulary entry. Many models tie it to the embedding table by reusing its transpose.
See it visualized- Autoregressive
Generating one token at a time, each conditioned on all the tokens before it, including ones the model itself just produced.
See it visualized- KV cache
The stored keys and values of earlier tokens, kept so that each new token only needs its own computed. It is valid because causal masking means old keys and values never change.
See it visualized- Cross-entropy loss
The standard training loss for next-token prediction: the negative log of the probability the model gave to the token that actually came next.
See it visualized- Gradient descent
Repeatedly nudging every parameter a small step against its gradient, so the loss goes down.
See it visualized