CHAPTER 12 · Observability, Hooks, and the Agent SDK
Observability, Hooks, and the Agent SDK
Everything so far has been the agent deciding what to do. This chapter is about the parts you control deterministically, regardless of what the model decides: hooks (code that fires on lifecycle events), budgets (hard limits on turns and cost), and the Agent SDK (the production loop exposed for programmatic use). This is the "observability" layer (layer 6), and its defining property is right there in the name: it is deterministic, not model-chosen.
Hooks: deterministic control points
A hook is a handler that fires at a fixed point in the agent's lifecycle. The contrast with skills is the key insight from the Claude Code architecture breakdown: a skill is invoked when the model decides it is relevant, but a hook fires every time its event occurs, no matter what the model wants. Skills are suggestions the model may take; hooks are rules the harness enforces.
The standard lifecycle events:
| Hook | Fires | Typical use |
|---|---|---|
PreToolUse | Before a tool runs | Block destructive commands; enforce policy |
PostToolUse | After a tool runs | Auto-format edited files; log the result |
Stop | When the agent finishes a turn | Aggregate results, notify |
SubagentStop | When a subagent finishes | Collect parallel results |
PreCompact | Before compaction runs | Archive the full transcript before it is summarized |
Handlers can be shell commands, HTTP webhooks, MCP tools, prompt judges, or (experimentally) agent-based verifiers. Two properties make hooks valuable: they are deterministic (a PreToolUse hook that blocks git push blocks it always, not "usually"), and they run in your process without consuming model context. So a hook is free in tokens, unlike an instruction you would otherwise stuff into the prompt and pay for every turn.
Notice how PreToolUse overlaps with Chapter 7's approval policy: both can block a tool before it runs. The difference is that the approval policy is a built-in classifier and a hook is your custom code. In practice you use both, the policy for the common safety tiers and hooks for project-specific rules.
Observability: knowing what happened
Beyond hooks, the harness records what it did. Claude Code writes every message, tool use, and result to a plaintext JSONL file under ~/.claude/projects/, which is what enables rewinding, resuming, and forking sessions (Chapter 2's session vocabulary). That log is your audit trail and your debugger. When an agent does something surprising, the JSONL is where you go to see the exact sequence of tool calls and results. For a system that runs code and holds credentials, an audit trail is not optional; Chapter 16's security incidents are partly stories about organizations that could not answer "what did the agent actually touch?"
The Agent SDK and budgets: making cost an engineering decision
The same production loop is exposed programmatically through the Agent SDK, for CI, services, and custom UIs. What matters for this chapter is the set of knobs it gives you, because they turn the vague worry "what if it runs forever?" into concrete configuration:
allowed_tools: which tools the agent may use.max_turns: cap the number of laps (the real version of Chapter 2's crude limit).max_budget_usd: a hard spending cap.effort: how hard the model should think.setting_sources: load the project'sCLAUDE.md, skills, and hooks.
And the result it returns carries a subtype telling you how it ended: success, error_max_turns, error_max_budget_usd, and so on, plus the per-session cost. The architecture breakdown's line is worth quoting: this makes "budget-aware agents an engineering task, not a hope." You do not cross your fingers that the agent stops; you set max_budget_usd and it stops. This directly addresses the Chapter 4 cost curve and the Chapter 16 warning about retry loops quietly running up a bill.
Here is the minimal SDK shape from the breakdown, so you recognize it:
import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions
async def run():
async for message in query(
prompt="Fix failing auth tests",
options=ClaudeAgentOptions(
allowed_tools=["Read", "Edit", "Bash", "Grep", "Glob"],
setting_sources=["project"], # load project CLAUDE.md, skills, hooks
max_turns=30,
),
):
... # handle each message; check the final ResultMessage.subtype