Skip to slide
Chapter 6 · Tools and Tool Execution
36 / 142

CHAPTER 06 · Tools and Tool Execution

Tools and Tool Execution

Tools are what make an agent agentic. Without them, the model can only talk. With them, it can read your code, edit files, run commands, search the web, and call external services. Anthropic groups the built-in tools into five families: file operations, search, execution, web, and code intelligence. But the families matter less than the mechanism: a tool is just a function the model can ask the harness to run, described to the model by a JSON schema (Chapter 3), and executed by the harness (never by the model itself).

In Chapter 2 we ran a tool with a single line: output = TOOLS[name](**args). That is fine for a demo and dangerous for anything real. A production tool call goes through a pipeline. The "Inside the Agent Harness" analysis lays out the six steps the harness performs when the model returns a tool call:

  1. Parse the arguments. The model returns JSON; the harness must parse it.
  2. Validate against the schema. Does the JSON match what the tool expects? Reject or repair if not.
  3. Check permissions. Is this allowed, or does it need approval? (Chapter 7.)
  4. Execute. Run it, ideally in a sandbox (Chapter 7).
  5. Capture output. Standard output, standard error, exit code, and timing.
  6. Truncate if needed, then format the result so the model can read success or failure clearly.

Steps 1, 2, 5, and 6 are this chapter. Steps 3 and 4's safety half is Chapter 7.

Structured outputs, not raw stdout

A theme worth hammering: the model needs to understand what happened, not just see a wall of text. Raw stdout is not enough. The "Inside the Agent Harness" piece is explicit that structured tool outputs are critical: you want the exit code, the timing, a truncation indicator, and clear success or failure signals. A command that prints nothing and exits 0 succeeded; a command that prints nothing and exits 1 failed. If you only forward stdout, the model cannot tell those apart. So wrap results in a small structure.

Here is the shape Codex uses for a long command's result:

Exit code: 0
Wall time: 1.23 seconds
Total output lines: 5000
Output:
[first 100 lines...]
... (4800 lines omitted) ...
[last 100 lines...]

Truncation: keep the head and the tail

Tool outputs can be enormous, and a 5,000-line log would blow the context budget (Chapter 4) all by itself. The clever move, which Codex calls token-aware truncation, is to preserve the beginning and the end while eliding the middle. Why those ends? Because the start of a log usually shows what was attempted and the end usually shows the result or the error. The middle is repetitive noise. A real-life analogy: if a friend sends you a 40-minute voice memo, you mostly want the first sentence ("here's what I tried") and the last ("and here's what broke"). The 38 minutes in the middle rarely change your reply.

Tools change with context

The available tools are not fixed. The harness adds and removes them based on permissions, sandbox mode, and what is present. Codex only exposes a view_image tool when an image is in play. Claude Code defers MCP tool schemas and loads them on demand (Chapter 10) so idle servers do not eat context. The takeaway for a builder: keep the tool list as small and relevant as you can, both to help the model choose well and to protect your cache (changing the tool list mid-session is a cache-buster, per Chapter 5).

← → arrow keys work too