20
Training stage

Modern LLM System

Production assistants wrap a model with context, memory, retrieval, and tools.

Every chapter up to this one has been a piece. This is where they sit inside something closer to a real product: a chat assistant, a coding tool, anything built around a model rather than being only the model.

Everything from Text through Transformer Stack, running as one forward pass per generated token.

This diagram is a general shape, assembled from publicly discussed concepts (routing between models, retrieval, tool use, memory), not a description of any specific product's actual architecture. ChatGPT, Claude, Gemini, and similar systems are built around foundation models, but their exact internal design, infrastructure, and routing logic are proprietary. Treat any specific claim about how one of them works internally with real skepticism unless it's something the company has documented itself.

An assistant is a model plus scaffolding

The model takes tokens and returns a distribution over the next token, and that is all it does. A product that feels like an assistant wraps that loop in software that decides what goes into the context and what happens to the output. Before each model call the application assembles the context: a system prompt with rules and persona, the conversation so far, anything retrieved, results from tools, and the latest user message. After the call it may filter, format or check the output before you see it.

Memory is text

The model has no memory between calls. What looks like memory is the application re-sending earlier messages, or saving notes somewhere and inserting the relevant ones into the next context. Either way the model sees it as tokens inside its window, which is why a long conversation eventually has to be trimmed or summarized.

Where each chapter lives

chapters 2-3     text and tokenizer (before the model)
chapters 4-11    the network: embeddings to the transformer stack
chapters 12-15   scores, softmax, sampling, the generation loop
chapters 16-17   training and fine-tuning (done before you use it)
chapters 18-19   retrieval and tools (around the model)

What a model alone doesn't provide

The weights hold what training captured, up to a cutoff date. Recent facts, private documents, exact arithmetic and anything that requires acting in the world come from the scaffolding: retrieval, tools and ordinary code. Building a reliable system is therefore a design problem as much as a modeling one. It decides what the model sees, limits what it may do, and checks what comes out.

What this lab leaves out

Models vary. Some use mixture-of-experts feed-forward layers, different attention layouts, other positional schemes or normalization, extra inputs such as images, and tricks for very long contexts. The lab shows one common decoder-only layout so the core mechanism is clear, and the chapters point to the main variants where they differ. Underneath all of them the same loop runs: context in, distribution out, sample, append.

The whole journey

A sentence becomes tokens, tokens become vectors, position gets folded in, attention lets each token gather information from the ones around it, a stack of transformer blocks refines that representation, a final projection turns it into scores over the whole vocabulary, softmax turns those scores into a distribution, and sampling picks one token, which becomes part of the context for doing all of it again. Training is what shaped every weight involved in that process in the first place. Everything on this page is that same loop, wrapped in more scaffolding.

Back to the beginning traces that whole path in one view.

Key Takeaway

A modern LLM product wraps a model with context assembly, routing, retrieval, memory, and tools, but the model at its center is still doing exactly what the earlier chapters described: predicting one next token at a time.