Glossary

Every term used across the lab, gathered in one place and linked back to where it's visualized.

Token

A single unit of text a model operates on: often a word piece, sometimes a whole word or a punctuation mark.

See it visualized
Tokenizer

The algorithm that splits text into tokens and maps each one to an integer id.

See it visualized
Embedding

A learned vector representation of a token, positioned so that related tokens end up near each other.

See it visualized
Context

The full sequence of tokens a model can see when producing its next prediction.

Attention

A mechanism that lets each token's representation incorporate information from other tokens, weighted by relevance.

See it visualized
Query

One of three projections used in attention. It represents what a token is 'looking for' in other tokens.

See it visualized
Key

One of three projections used in attention. It represents what a token 'offers' when compared against a query.

See it visualized
Value

One of three projections used in attention. It carries the information that actually gets passed on, weighted by attention score.

See it visualized
Logit

A raw, unnormalized score the model assigns to a candidate next token, before softmax.

See it visualized
Softmax

A function that turns a list of raw scores into a probability distribution that sums to 1.

See it visualized
Temperature

A parameter that reshapes a probability distribution before sampling: lower values sharpen it, higher values flatten it.

See it visualized
Transformer

The neural network architecture built from alternating attention and feed-forward layers that underlies most modern language models.

See it visualized
Residual connection

A shortcut that adds a layer's input back to its output, which helps very deep networks train.

See it visualized
LayerNorm

A normalization step in a transformer block: it re-centers and rescales each token's vector using that vector's own mean and spread, keeping values in a stable numerical range.

See it visualized
Feed-forward network

A small per-token neural network applied after attention, transforming each token's representation independently.

See it visualized
Parameter

A learned numerical weight inside the model. Modern models have anywhere from millions to hundreds of billions of them.

Gradient

A measure of how much a small change in a parameter would change the loss, used to decide how to update that parameter during training.

See it visualized
Loss

A single number measuring how wrong the model's predictions were on a batch of examples during training.

See it visualized
Backpropagation

The algorithm that computes gradients for every parameter in the network from a single loss value.

Fine-tuning

Further training a pretrained base model, typically to make it follow instructions or match a particular style.

See it visualized
RAG

Retrieval-augmented generation: retrieving relevant documents and adding them to context before generating an answer.

See it visualized
Vector database

A database optimized for finding the nearest embedding vectors to a query vector, used to power retrieval.

Inference

Running a trained model to produce an output, as opposed to training it.

Sampling

Choosing the next token from a probability distribution, not always the single most likely one.

See it visualized
Causal mask

A rule inside attention that stops each token from seeing the tokens after it, so a text generator can't peek at the answer. It is applied by adding −∞ to the forbidden scores before softmax, which turns their weights into exactly zero.

See it visualized
Positional encoding

A signal that tells the model where each token sits in the sequence, since attention alone has no sense of order. It can be added to the embeddings (sinusoidal or learned) or applied inside attention (rotary embeddings, ALiBi).

See it visualized
Multi-head attention

Several attention computations run in parallel, each with its own query, key and value projections in a smaller subspace. Their outputs are concatenated and mixed by one more learned matrix.

See it visualized
Dot product

Multiply two vectors coordinate by coordinate and add up the results. The number is large when the vectors point in similar directions, which is why it serves as the similarity score in attention.

See it visualized
Cosine similarity

A similarity measure that depends only on the angle between two vectors: 1 for the same direction, 0 for unrelated, −1 for opposite. Commonly used to compare embeddings, including in retrieval.

See it visualized
Vocabulary

The fixed set of tokens a model can read and write, each with an integer ID. Sizes range from tens of thousands to a few hundred thousand entries.

See it visualized
Hidden state

A token's vector at some point inside the network, between the embedding and the output scores. The one after the last block is what gets projected into logits.

See it visualized
Residual stream

The running vector each token carries through the stack. Every attention and feed-forward sub-layer reads it and adds its result back, so the final vector is the original embedding plus everything added along the way.

See it visualized
Unembedding

The final matrix that turns a token's last vector into one score per vocabulary entry. Many models tie it to the embedding table by reusing its transpose.

See it visualized
Autoregressive

Generating one token at a time, each conditioned on all the tokens before it, including ones the model itself just produced.

See it visualized
KV cache

The stored keys and values of earlier tokens, kept so that each new token only needs its own computed. It is valid because causal masking means old keys and values never change.

See it visualized
Cross-entropy loss

The standard training loss for next-token prediction: the negative log of the probability the model gave to the token that actually came next.

See it visualized
Gradient descent

Repeatedly nudging every parameter a small step against its gradient, so the loss goes down.

See it visualized