Tools
The model can request an external action instead of answering directly from its parameters.
Nothing covered so far lets a model check the current weather, run code, or look anything up; it can only produce tokens based on its parameters and its context. Tool calling closes that gap by giving the model a way to ask for something to be done on its behalf, mid-generation.
User
What's the weather in Austin right now?
User message
The division of labor is the whole idea: the model's job stops at generating a well-formed request. It doesn't call the weather API itself. It can't: it has no network access of its own. Whatever system is running the model is what actually executes the tool and feeds the result back in as new context.
A tool call is just text
A language model can only output tokens. A "tool call" is a particular text format that the model has been fine-tuned to produce when it needs something it can't do alone, like the JSON request in the demo. The system around the model spots that text, runs the real function, and puts the result back into the context as more text. The model then continues with the answer in front of it. The model never runs anything itself.
The loop
1. the prompt describes each tool: name, purpose, arguments
2. the model writes a call: get_weather("Austin, TX")
3. the system runs it and gets: 84°F, partly cloudy
4. the result is appended to the context
5. the model writes the final answer using that resultHow the model knows when to call
Nothing special happens inside the network. The tool descriptions are part of the prompt, and fine-tuning taught the model that for some requests the most probable continuation is a call in the expected format, which the system can detect and stop generation at. Whether to call, which tool, and with what arguments are all next-token predictions like any other, so they can be wrong: the wrong tool, a missing argument, a call when none was needed.
Agents: the loop, repeated
An agent is this loop run several times. The model calls a tool, reads the result, decides what to do next, and repeats until it has an answer or hits a limit. Each round adds text to the context, so long tasks use up the window, and a small error early on can compound. Systems normally cap the number of steps and validate arguments before running anything.
Trust and safety
Tool results come back as text in the context, and a model can't reliably tell instructions from data. A web page or document containing a sentence such as "ignore your previous instructions and..." can try to steer the model, which is known as prompt injection. So it is the system, not the model, that has to enforce permissions: which tools exist, what they can touch, and which actions need a person's confirmation.
Tools, retrieval, memory and the system prompt are all ways of putting the right text into the context. The last chapter looks at how a complete assistant assembles them.