
Part 3 of the AI reading series. Read this after Part 1: Where Machine Learning Came From and Part 2: How Connectionism Conquered Language.
By now you know what a Transformer is: an attention machine that, trained to predict the next word over oceans of text, becomes a language model. Time to look this machine in the eye and ask the uncomfortable question: what can it NOT do?
Try the thought experiment. Ask an LLM:
“What’s the weather in Chiang Mai right now?” — It cannot know. Its knowledge stops at the end of its training data (its knowledge cutoff), and it has no access to the outside world.
“What is 847 × 3,921?” — It will answer confidently… and often wrongly. An LLM does not calculate: it predicts plausible tokens. Multiplication is not text prediction.
“Book me a table for tonight.” — It will produce a polite, perfectly useless reply: an LLM generates text, it does not act.
And when it doesn’t know, it does worse than stay silent: it hallucinates — inventing facts, dates, and references with the same aplomb as when it is right. This is not a bug but a direct consequence of its training objective: produce the most probable continuation of text, not the most true. In the
terms of the first post in this series, the horizon of the computation is the plausibility of the next token, nothing more.
The four walls of the box:
1. Frozen knowledge — the model is a snapshot of the past.
2. No reliable calculation — arithmetic, exact logic, and executed code escape it.
3. No access to the world — neither reading (the web, your files, a database) nor writing (sending, booking, modifying).
4. Hallucinations — confidence does not guarantee truth.
The idea at the heart of this post fits in one sentence: rather than trying to cram all the world’s knowledge and skills into the model’s weights, give it tools — a calculator, a search engine, a code interpreter, a weather API — and teach it to use them. Just like humans, whose intelligence lives not only in their brains but in their ability to use a pencil, a book, or a computer.
1. But how can a text generator “use” anything?
This is the central question, and its answer is disarmingly elegant: since the LLM can only produce text, let's make tool use… text.
The mechanism, in principle, takes four steps:
1. Describe the tools to the model, in text. In the prompt (typically the system prompt, the hidden instruction that precedes the conversation), you write something like: "You have a tool weather(city) that returns current weather, and a tool calculator(expression) that evaluates a math expression. To use a tool, write name(arguments)."
2. The model generates a call instead of an answer. Faced with "What's the weather in Chiang Mai?", the model — if it has learned to — generates weather("Chiang Mai") and stops.
3. Ordinary code intercepts that text and actually runs the tool. This is the crucial point: the LLM executes nothing itself. A plain piece of code (roughly, a Python loop) detects the call, queries the real weather API, and gets the result: .
4. The result is injected back into the context, as if it were the next part of the conversation, and the model resumes generating: "It's currently 31°C in Chiang Mai, with thunderstorms…"
Reread those four steps: at no point did we modify the Transformer's architecture. Everything happens in the text — tool descriptions, call, result. This is what we now call function calling, and the loop that executes things is called the runtime or orchestrator. Three remarks worth keeping, because they come back constantly:
- The model decides; the runtime executes. This separation is fundamental, including for safety: you can refuse to execute a call, ask a human for confirmation, or restrict the available tools (sandboxing).
- Everything lives in the context. Tool descriptions, calls, results — all of it must fit in the model's context window. Now you know why the size of that window (4k tokens yesterday, hundreds of thousands today) is a major industrial battleground.
- Format matters enormously. In practice, calls are emitted in JSON (a standard text format for structured data: ) and tools are described by a schema (name, description, expected parameters and their types). The clearer the description, the better the model uses the tool — it is prompt engineering applied to tools.
2. A short history of tool-using LLMs (2020–2025)
Here is how it actually took hold.
2020–2021 — The first hacks. As early as GPT-3, researchers noticed that answers improve when
you inject relevant documents, found by a classic search engine, into the prompt. This is the birth of
RAG (Retrieval-Augmented Generation, Lewis et al., 2020) — technically a cousin of function
calling: the context is augmented with external information, but it is the system, not the model, that
decides what to fetch. In late 2021, OpenAI unveils WebGPT: a GPT-3 fine-tuned to navigate a text
based browser (search, click, cite sources) to answer questions. The first demonstration that an
LLM can drive a complex tool.
February 2023 — Toolformer, the paper you’re about to read. Meta AI (Schick et al.) frames
the question differently: can a model learn to use tools without massive human demonstrations?
Their answer is a model that teaches itself where to insert API calls into text, keeping only the calls
that genuinely improve its predictions. It is a pivotal paper — simple in its idea and deeply revealing
of the connectionist method: even tool use is not programmed, it is learned.
March–June 2023 — Industrialization. OpenAI launches ChatGPT plugins (March), then exposes
function calling directly in its API (June): developers provide the list of tools as JSON Schema, and
the model answers with structured calls. Anthropic, Google, and open models (Llama, Mistral,
Qwen) follow with the same de facto standard. Tool calling stops being a research hack and
becomes a core feature, explicitly trained into models — the chat templates of models on Hugging
Face now include special tokens dedicated to tool calls.
2024–2025 — Standardization. With every vendor describing tools its own way, the ecosystem
fragments. In late 2024, Anthropic publishes MCP (Model Context Protocol), an open protocol for
connecting any data source or service to any model — a kind of “universal USB port” for tools,
widely adopted in 2025. At the same time, tools become spectacular: code execution in sandboxes,
full computer control (mouse, keyboard, screenshots — computer use), web browsing… The line
between “answering” and “acting” blurs. That is exactly the subject of the next post in this series,
on agents.
3. Rereading all this through Cardon’s grid
Take a minute to place this evolution in the world / calculator / horizon grid from the first post.
- The world of the pure LLM was its training corpus — a frozen world, absorbed once and for all into the weights. With tools, the world comes alive again: the model queries reality at the moment of the question. Three quarters of a century later, we are back to Wiener’s cybernetic loop: perceive, act, measure the gap, correct.
- The calculator becomes hybrid. The neural network (intuitive, statistical, fallible) delegates to classical programs (exact, deterministic, verifiable) what they do better: calculating, searching, executing. Some see it as a quiet revenge of symbolic AI: the calculator, the Python interpreter, the database are rule machines — and the network learns when to invoke them. The two paradigms of Cardon’s paper, reconciled in practice.
- The horizon widens: from “produce the most plausible text” to “accomplish the requested task,” relying on verifiable tool results. Hallucination does not disappear, but it recedes wherever a tool can ground the answer in a source.
4. How to read “Toolformer”
The paper “Toolformer: Language Models Can Teach Themselves to Use Tools” (Schick et al., 2023, about 20 pages with appendices) is very readable. Roadmap:
1. Abstract + Introduction: the general idea and the five tools used (calculator, question answering, search, translation, calendar). Note their deliberate simplicity.
2. Section 2 (the method — the heart of the paper): read it twice. The self-teaching mechanism takes three steps: (a) the model itself, prompted with a few examples, annotates texts with candidate API calls; (b) those calls are executed; (c) only the calls whose result reduces perplexity — that is, genuinely helps the model predict what comes next — are kept. The model is then fine-tuned on these filtered, annotated texts.
3. Figure 1 and the examples: concrete, and worth a thousand equations.
4. Sections 3–4 (experiments): skim. Remember the headline result: a 6.7-billion-parameter model equipped with tools beats GPT-3 (175 billion) on several tasks. Tools can substitute for scale — a point worth pondering.
5. Limitations (end of the paper + your own critical eye): Toolformer cannot chain tools, nor use the result of one call to decide on another. Hold on to that frustration: it is precisely what leads to agents.
Three questions to carry through the paper: Why is the criterion “does the call reduce perplexity?” such a good idea? What does it
guarantee — and what does it not? How is Toolformer’s method faithful to the “inductive move” described by Cardon (empty the calculator, let the data decide)? What’s the difference between modern function calling (section 1 of this post) and Toolformer’s approach? (Hint: training vs. prompting.)
Timeline
| Year | Event | Why it matters |
|---|---|---|
| 2020 | RAG (Lewis et al.) | Augmenting the context with retrieved documents |
| 2021 | WebGPT (OpenAI) | An LLM fine-tuned to drive a browser |
| 2022 | Chain-of-thought; ReAct | Reasoning step by step; reasoning AND acting (next post) |
| 2023 | Toolformer (Schick et al., Meta) |
The model teaches itself to use tools |
| 2023 | ChatGPT plugins / Function calling in the OpenAI API |
Tools go mainstream / Tool calling becomes an industry standard |
| 2024 | Native function calling everywhere (Claude, Gemini, Llama, Qwen…) |
Special tokens, dedicated training |
| 2024 | MCP (Anthropic) | An open, universal protocol for tool connections |
| 2025 | Code execution, computer use, browsing |
From “answering” to “acting” — heading for agents |
Glossary
API (Application Programming Interface) — A service's standardized entry point: you send a formatted request ("weather in Chiang Mai?"), you receive a structured response. LLM tools are almost always APIs.
Context window — Everything the model "sees" when generating: system prompt, conversation, tool descriptions, call results. A limited and precious resource.
Few-shot / zero-shot — Using the model with a few examples in the prompt (few-shot) or none (zero-shot). Toolformer uses few-shot prompting to bootstrap its self-annotation.
Function calling — An LLM's ability to emit, as structured text, a request to execute a tool described in its context. The model asks; the runtime executes.
Grounding — Anchoring the model's answers in verifiable sources (tool results, documents, calculations) rather than in its statistical memory alone. A partial antidote to hallucination.
Hallucination — Producing false information with the confidence of truth. A structural consequence of the "predict the plausible token" objective.
JSON — The universal text format for structured data (``). The lingua franca of tool calls.
JSON Schema — A formal description of a tool: its name, its parameters, their types. What the model reads to know how to call it.
Knowledge cutoff — The end date of the training corpus. Beyond it, the model knows nothing — unless it has a search tool.
MCP (Model Context Protocol) — An open protocol (2024) standardizing connections between models and tools/data sources. The "USB port for tools."
Orchestrator / runtime — The ordinary program that runs the loop: send the context to the model, detect calls, execute tools, inject results back.
Perplexity — A measure of the model's "surprise" at a text: low = the model predicts well. The central filtering criterion in Toolformer.
RAG (Retrieval-Augmented Generation) — An architecture that searches for relevant documents (in a database, on the web) and injects them into the context before generating an answer.
Sandbox — An isolated, restricted execution environment where the model's actions (code, for instance) run without risk to the host system.
Special token — A reserved vocabulary token (invisible in normal text) used to mark structure: start of a tool call, end of a message, and so on. Recent models have them for function calling.
System prompt — An instruction invisible to the user, placed before the conversation, that defines the model's role and, often, its tools.
Watch After Reading
- “[1hr Talk] Intro to Large Language Models” — Andrej Karpathy: his “LLM OS” vision places tool use (browser, calculator, Python interpreter) in a memorable big picture.
- “What is Function Calling?” and “What is RAG?” — IBM Technology (~8 min each): two short formats to anchor the industry vocabulary.
Happy reading — and while you read, keep Toolformer’s limitation in mind: one call at a time, no chaining. The logical next step awaits in the next post: what if we let the model loop?
Table of contents

Leave A Comment