18
Training stage

RAG

Retrieved documents are added to context so the model can answer with information it wasn't trained on.

Query→Query embedding→Vector search→Top-K documents→Context→ answer

Everything a model "knows" from training is baked into its parameters, and changing it means training again, which is slow and expensive. Retrieval-augmented generation sidesteps that: find documents relevant to the query using embedding similarity, and put them directly into context, where the model can read them like anything else you'd type in.

Query: “What's the return policy for electronics?”

Top-K documents retrievedK = 3

Electronics may be returned within 30 days of purchase with the original receipt and packaging.

0.91

Opened software and digital downloads are final sale and cannot be returned.

0.74

A 15% restocking fee applies to returned electronics missing their original box.

0.69

Extended warranties for electronics can be purchased at checkout or within 90 days after.

0.52

Clothing and footwear can be returned within 60 days if unworn with tags attached.

0.31

Curbside pickup orders are held for 5 business days before being restocked.

0.17

Gift cards do not expire and are not redeemable for cash except where required by law.

0.11

Store hours are 9am to 8pm Monday through Saturday, 10am to 6pm on Sunday.

0.08

Context sent to the model

The query, plus 3 retrieved chunks (292 characters total) are placed in context together, and the model generates an answer grounded in them.

Similarity scores here are hand-assigned to demonstrate ranking; a real system computes them from the cosine similarity between the query's embedding and each chunk's embedding.

Parameter knowledge vs. retrieved context

These are genuinely different sources of information to the model. Parameter knowledge is compressed, general, and fixed at training time. It can't cite a specific document, and it can be wrong in ways that are hard to predict. Retrieved context is specific, current as of retrieval, and directly inspectable, but only as good as what the search step finds, and a large enough K can push relevant context out of the window entirely.

A higher K isn't automatically better. More retrieved chunks means more context for the model to search through, and to potentially get distracted or confused by, not just more available information.

What retrieval adds

RAG leaves the model alone and changes what it reads. Before the model answers, the system looks up passages likely to help, pastes them into the prompt, and lets the model answer with that text in front of it. The model uses the passages the way it uses any context: attention can read from them exactly as it reads from the question.

The pipeline

documents   split into chunks, each turned into a vector
            (done once, ahead of time)
question    turned into a vector by the same embedding model
search      chunks whose vectors are closest to the question's
prompt      retrieved chunks + the question
model       writes the answer with those chunks in context

Retrieval is embeddings again

The search step reuses an idea from chapter 4. An embedding model, often a separate and smaller network trained so that texts with similar meaning get similar vectors, maps each chunk to a vector ahead of time and stores the vectors in an index. At question time the question is embedded the same way, and the chunks whose vectors are most similar come back, usually measured by cosine similarity:

score(q,c)=eq⋅ec∥eq∥ ∥ec∥\text{score}(q, c) = \dfrac{e_q \cdot e_c}{\lVert e_q \rVert \, \lVert e_c \rVert}

With millions of chunks, the index uses approximate nearest-neighbour search instead of comparing the question with every vector. The similarity scores in the demo are hand-assigned so the ranking is easy to read; in a real system they come from this formula. The slider changes KK, how many of the top-ranked chunks get pasted in.

The trade-offs

Everything retrieved has to fit in the context window, and every token in it costs time and money, so KK is a balance. With too few chunks the answer may not be among them. With too many, the relevant passage is surrounded by noise, and models have been reported to use information at the start and end of a long context better than information in the middle. Chunking matters too: a fact that is split across two chunks may be retrieved only in part.

Where RAG fails

If retrieval misses the right passage, the model has nothing to go on and may answer from memory or invent something, often in the same confident tone as a correct answer. If the retrieved text is wrong, outdated or contradictory, the model can repeat it faithfully. RAG lowers the chance of unsupported answers and makes answers easier to check, since the sources can be shown, but it doesn't eliminate mistakes. The quality of the retrieval step is usually the main limit on the quality of the answer.

RAG decides what the model is shown, using a fixed lookup. Tool use goes a step further: the model itself decides what to ask for.

Key Takeaway

RAG retrieves documents relevant to a query using embedding similarity and places them in context, letting a model answer using specific, current information it was never trained on.