Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

What is tool use (a.k.a. function calling)?

A model that only emits text somehow ends up booking your flight. The trick isn't in the weights — it's in the contract between model, harness, and your code.

AI & ML intro May 7, 2026 · updated Aug 25, 2026 · 13 min read

On this page

The picture version

Six pictures for a reader who has never seen the word “tool” used this way. The prose below fills in the seams the pictures skip.

1 · The problem

It has no clock, no thermometer, and no network. It answers anyway.

“what’s the weather in Tokyo right now?” THE MODEL weights frozen long ago no clock no thermometer no network tokens in, tokens out. that is all. “It’s 14°C and lightly raining in Tokyo.” and it is actually correct So the number did not come out of the model. something else fetched it, and the model was handed the answer before it wrote that sentence
A raw model can think about the world but cannot touch it. Tool use is the seam where text prediction meets the rest of your software — and the pictures below build it one failure at a time.

2 · The naive attempt

Just ask it. It invents a temperature, in exactly the same tone.

what it would need to say “I don’t have this.” no token sequence represents an absence of data what it says instead “It’s 18°C and clear.” the most plausible-looking continuation, which is a number There is no failure signal. The wrong answer looks like the right one. so the fix cannot be “ask more nicely” — it has to be a different arrangement around the model
The model produces the most likely continuation, and for “what is the temperature in Tokyo” that is a temperature. Nothing in the output distinguishes a fetched fact from an invented one, which is why the rest of the machinery exists.

3 · Two fixes, stacked

Hand it a menu — then train it to order from the menu.

fix 1: put the menu in the prompt name: get_weather what it does: current weather for a city args: city, units but a base model writes… “Great question! To use the get_weather function, you would first want to…” a tutorial. it has read far more text about APIs than calls to them. fix 2: post-train it on thousands of examples of the contract “given this menu + this question, the right next thing to emit is…” {"name": "get_weather", …} tool use is not in the architecture — it is a behaviour the model was trained to perform when the prompt has this shape
The menu alone is not enough, because emitting a call is not the most likely continuation for a model trained on the open internet. Post-training is what shifts the probability onto the call itself — which is also why the model can still pick the wrong tool, or invent one.

4 · One unlucky token

A missing brace and the whole call is garbage.

still sampling one token at a time, so: {"name": "get_weather", "args": {"city": "Tokyo" ← unclosed not JSON any more — the host has nothing to run, and a feature that works 99 times in 100 is not something to guarantee the fix: at every step, delete the tokens that could not keep it valid what the model wanted "Tokyo" "Osaka" Sure! <newline> probability set to zero the model still chooses the city it just cannot go off the rails of the grammar OpenAI documents this mechanism for its strict Structured Outputs mode; other providers promise conformance without always naming the mechanism
Constrained decoding is what turns “usually parses” into something a provider can put a guarantee on. It prevents the parse error, never the semantic one — a perfectly valid call to the wrong tool sails straight through.

5 · The loop

The model slides a note under the door. Someone else opens it.

THE MODEL emits text. nothing else. THE HARNESS your code. the one with hands. the real weather API over the actual network 1 get_weather (Tokyo) validates the args against the schema asks you to approve — if its policy says to 2 3 14°C, light rain 4 pasted into the context “It’s 14°C in Tokyo.”   One continuous-looking reply. Two model calls and a real HTTP request.
The model’s part is over the instant the call is emitted; everything after that is the harness. The model picks, the harness enforces — a delete_everything tool will be called if the model thinks it should be, and the only thing between that thought and the action is your permission layer.

6 · Keep this card

The whole thing on one index card.

tool use = the model emits a structured call + the host executes it + the result goes back into the context ∴ “does this model support tool use?” is a question about a whole stack
Picture to keep: a brilliant consultant locked in a room with no phone, who can only slide written requests under the door and wait for someone outside to slide the answer back. Where the picture breaks: the slip of paper has to match a machine-checked form exactly, and the person outside is under no obligation to do what it says.

Why it exists

Ask ChatGPT or Claude “what’s the weather in Tokyo right now?” and you’ll often get back an actual current temperature. Pause on that for a second. The model’s weights were frozen long before today. It has no clock, no thermometer, and no internet connection wired into the matrix multiplications. All it does — at the lowest level — is map a sequence of tokens to a probability distribution over the next token. So how did the right number end up on your screen?

The short version: the model didn’t fetch the weather. It emitted some text of a particular shape — something like {"name": "get_weather", "args": {"city": "Tokyo"}} — and a different program, sitting between you and the model, recognized that shape, called the real weather API, and pasted the result back into the conversation before asking the model to continue.

That dance is “tool use” (when the docs are talking about agents) or “function calling” (when they’re talking about an API). It exists because a raw LLM can think about the world but can’t touch it. Tool use is the seam where text-prediction meets the rest of your software. Without it, every assistant is a parlor trick that ends at the edge of its training data.

Why it matters now

The AI products that do anything beyond talking are tool-using ones. Coding agents that read your files and run your tests, chatbots that search the web, assistants that book travel or file tickets, anything wired into MCP — under the hood, all of them are doing the same loop: model emits a structured call, host runs it, result goes back into context.

This is also where a large share of agent failures land. The model picks the wrong tool, hallucinates an argument that looks plausible but doesn’t exist, calls the same tool ten times in a row, or misreads the result and confidently lies about it. None of those are “the model isn’t smart enough” — they’re specific failure modes of the tool-use contract. You can’t reason about why your agent broke without a clear picture of what tool use actually is.

It’s also the layer where the API providers compete most directly. OpenAI, Anthropic, Google, and the open-model serving stacks each ship slightly different tool-use surfaces, but the shape they’ve landed on is recognizably the same: a JSON schema for each tool, and a way for the model to emit a call against that schema. How strongly each one promises the call will parse — and by what mechanism — varies, which is the subject of a caveat further down.

The short answer

tool use = the model emits a structured call + the host executes it + the result is fed back into the context

Picture to keep: a brilliant consultant locked in a room with no phone, who can only slide written requests under the door and wait for someone outside to slide the answer back. Where the picture breaks: the slip of paper isn’t a note in English — it has to match a machine-checked form exactly, and the person outside is under no obligation to do what it says.

The model never runs anything. It just produces text in a format the host agrees to interpret as “please run this function.” The host runs it, appends the output to the conversation, and asks the model to continue. Loop until the model decides it’s done.

How it works

Build it yourself, starting from the obvious thing, and each piece shows up as the fix for the previous piece’s failure.

Naive attempt: just ask the model. → It makes the weather up.

Send “what’s the weather in Tokyo?” to a raw model and you get a fluent, confidently-worded temperature that came from nowhere. There’s no failure signal — the model has no way to represent “I don’t have this,” so it produces the most likely-looking continuation, which is a plausible number.

Fix: put the tool list in the prompt

When you start a conversation that supports tools, the host sends the model a list of available tools. Each tool has a name, a one-line description, and a JSON schema for its arguments — exactly the kind of metadata you’d write into an OpenAPI spec. In a typical API call:

{
  "tools": [
    {
      "name": "get_weather",
      "description": "Get the current weather for a city.",
      "input_schema": {
        "type": "object",
        "properties": {
          "city":  {"type": "string"},
          "units": {"type": "string", "enum": ["celsius", "fahrenheit"]}
        },
        "required": ["city"]
      }
    }
  ]
}

The provider’s SDK formats this into whatever the model was trained to read — often a section of the system prompt, sometimes a dedicated channel in a chat template. The model sees the tools the same way it sees any other text: as tokens.

So when people say “the model decided to call get_weather,” that’s shorthand. What actually happened is: the model conditioned on (your question + the description of get_weather) and the next-token distribution put high probability on a sequence of tokens that, decoded, spells out a tool call.

But a base model reads that list and writes a tutorial

Describing the tool isn’t enough. A model trained only on internet text has seen far more text about weather APIs than text that is a call to one, so “here is a get_weather function” plus “what’s the weather in Tokyo?” is just as likely to produce a helpful explanation of how you might call it.

Fix: post-train it on the contract

What makes function calling work is post-training — supervised fine-tuning, and generally some form of preference-based training such as RLHF — on examples of the form “given this tool list and this user question, the right next thing to emit is this tool call.” After enough examples, the model has internalized: “when a request needs information I don’t have, and a relevant tool is in the list, the right move is to emit a call against that tool’s schema, not to guess.”

This is the part that’s easy to miss. Tool use is not a property of the model architecture; it’s a behavior the model was trained to perform when the prompt has the right shape. The model can still ignore tools, call the wrong one, or invent a tool that doesn’t exist. Better post-training narrows those failure modes; it doesn’t eliminate them.

But one stray token still breaks the parse

Even with great post-training, the model is still sampling tokens one at a time. A single unlucky draw — a stray comma, a missing brace — and {"name": "get_weather", "args": {"city": "Tokyo"} isn’t JSON any more. The host has nothing to run, and a feature that works 99 times out of 100 isn’t something a provider can put a guarantee on.

Fix: constrain the decoder

The cleanest way around this is constrained decoding: once the model has decided it’s emitting a tool call, the inference server tracks which tokens can still keep the JSON valid given the prefix so far, sets the probability of every other token to zero, renormalizes, and samples from what’s left. The model still chooses the city name and the units, weighted by its own distribution — it just can’t go off the rails of the grammar.

OpenAI has documented this explicitly: their “Structured Outputs” mode (strict: true) uses constrained decoding to guarantee that the call matches the supplied JSON Schema. Other providers offer their own schema-conformance affordances for tool use, but they don’t always spell out the mechanism — constrained decoding, heavier post-training, and internal retries are all in use, and which of them sits behind a given provider’s promise is frequently not public. Check the specific promise your provider makes rather than assuming a universal one.

(For the longer version of this idea, see Why is structured output so hard? — function calling is structured output with a per-tool schema and a “when to call” prior baked in.)

And the model still can’t run it — so the host does

The model’s part is over the moment the tool call is emitted. From here on, the agent harness takes over:

user:     "what's the weather in Tokyo?"

→ model emits: tool_call(get_weather, {city: "Tokyo", units: "celsius"})
→ harness validates the args against the schema
→ harness asks the user to approve (only if that harness's policy says to)
→ harness calls the real weather API
→ harness appends to context: tool_result(get_weather, "14°C, light rain")
→ model emits: "It's 14°C and lightly raining in Tokyo right now."
→ harness exits the loop, returns the answer

The model never touched the network. It saw a question, emitted a structured call, then — on the next turn, with the result now in its context — produced the friendly answer. From the user’s perspective it felt like one continuous reply, but it was at least two model calls and a real HTTP request stitched together by the host.

Where it gets subtle

You started with tool use = the model emits a structured call + the host executes it + the result goes back into context. What did this post add? — + the emission has to be engineered into being reliable: the tool list has to be in the prompt, the model has to have been post-trained on the contract, and something has to keep the JSON valid. A function-calling model isn’t doing anything fundamentally different from a regular one — it’s still emitting the most likely next tokens. Those arrangements are what make the most likely next tokens spell out a call into your code, which is also why “does this model support tool use?” is a question about a whole stack rather than a weight file.

So: that Tokyo temperature never came out of the model. It came out of a weather API, and the model’s entire contribution was writing a well-formed request for it and then reading the answer back to you in a sentence.

Going deeper

Well-established, and in provider docs or published work: that the model itself doesn’t execute anything (the host does), that post-training is what teaches a model to emit calls in the right shape, and that OpenAI’s strict Structured Outputs mode is documented to use constrained decoding. What’s less publicly specified is the exact mechanism each provider uses for strict tool use; some are clearly constrained decoding, others may be a mix of constrained decoding, heavy fine-tuning, and internal retry. The exact per-provider wire format and which specific models support which tool-use features changes often; treat any specific claim about “model X supports parallel tool calls as of date Y” as something to verify against current docs rather than memorize.