17
Training stage

Fine-Tuning

A pretrained base model is further trained to follow instructions reliably.

Pretraining→Base model→Instruction data→Fine-tuning→Instruction-following model

The training chapter covered the core loop: predict the next token, measure loss, update weights. Pretraining runs that loop over huge amounts of raw text, with no notion of "questions" or "answers." The model just learns to predict plausible continuations, which on its own produces something that completes text, not something that follows instructions.

Base model completion

“Explain photosynthesis. Use at least 200 words and include a diagram. This assignment is due Friday and will be graded on accuracy and clarity...”

Trained only to continue text, it treats the prompt as the start of a document, plausibly a homework assignment, and keeps writing in that style.

Instruction-tuned response

“Photosynthesis is the process plants use to convert sunlight into chemical energy. Using sunlight, water, and carbon dioxide, plants produce glucose and release oxygen as a byproduct...”

Fine-tuned on examples of instructions paired with helpful responses, it has learned to treat the prompt as a request to fulfill, not text to continue.

Fine-tuning is more of the same training loop, applied to a much smaller, carefully curated dataset: instructions paired with the kind of response a helpful assistant should give. Nothing about the underlying mechanism changes; what changes is what the model is being trained to predict.

"Fine-tuning" covers a range of specific techniques in practice, including further supervised training on examples and reinforcement-learning-based methods that optimize for human preferences between responses rather than matching one fixed target output. The comparison above illustrates the general shift in behavior, not one specific method.

Same objective, different data

Fine-tuning continues training from the pretrained weights, with the same kind of loss, on a much smaller and more carefully chosen dataset. In supervised fine-tuning (SFT) the data are demonstrations: pairs of a prompt and the kind of response you want, often written or vetted by people. The loss is often computed only on the response tokens, so the model is graded on what it should answer and not on reproducing the question:

LSFT(θ)=−∑t ∈ responselog⁡Pθ(xt∣x<t)\mathcal{L}_{\text{SFT}}(\theta) = -\sum_{t \,\in\, \text{response}} \log P_\theta\big(x_t \mid x_{<t}\big)

The prompt is still fed in as context, so the gradient flows through everything the response depended on. Only the tokens being scored change.

From base model to assistant

A common sequence of stages

stage              data                     result
pretraining        huge raw text corpus     base model
supervised tuning  prompt + wanted reply    follows instructions
preference tuning  answers ranked by people matches preferences

A base model is an autocomplete engine. Given a question, it may continue with more questions, because that is what text often looks like. SFT teaches it the format of being asked and answering. A further stage, preference tuning, learns from comparisons rather than demonstrations. In reinforcement learning from human feedback (RLHF), a separate reward model is trained to predict which of two answers people prefer, and the language model is then adjusted to produce answers that score higher, with a penalty that keeps it from drifting too far from where it started. Direct preference optimization (DPO) gets a similar effect straight from the comparisons, without a separate reward model. Details vary a lot between labs, so read this as the general shape.

Changing less than everything

Full fine-tuning updates every weight, which is expensive. A popular shortcut, LoRA, freezes each original matrix WW and trains a small low-rank correction next to it:

W′=W+BA,B∈Rd×r,A∈Rr×d,r≪dW' = W + BA, \qquad B \in \mathbb{R}^{d \times r}, \quad A \in \mathbb{R}^{r \times d}, \quad r \ll d

With rr of a few dozen at most and dd in the thousands, the trainable part is a tiny fraction of the model, yet the update can still shift its behavior noticeably, and only the small matrices need to be stored for each task.

What fine-tuning is good at

It is strongest at changing behavior: tone, format, following a particular kind of instruction, declining certain requests. It tends to be a less reliable way to teach a model many new facts, because a modest number of examples has to reshape the weights, and for facts that change or need sourcing, retrieval (next chapter) is often the better tool. The two completions in the demo are hand-written to illustrate the difference in behavior. They are not outputs of real models.

A fine-tuned model runs exactly the same forward pass as the base model. What has changed is the numbers inside. The next chapters leave the weights alone and change what the model is given.

Key Takeaway

Fine-tuning doesn't introduce a new mechanism. It's the same predict-and-update loop, redirected at a curated dataset of instructions and good responses, and that redirection is what turns a raw text-completion model into something that behaves like an assistant.