CHAPTER 07 · Sandboxing, Approvals, and Checkpoints
Sandboxing, Approvals, and Checkpoints
An agent that can run arbitrary shell commands is, by definition, an agent that can do real damage: delete files, leak secrets, hit the network, push to production. Safety is not a nice-to-have here; the "Inside the Agent Harness" analysis calls it "table stakes." This chapter covers the three controls that make running a tool-using agent sane: sandboxing (limit what a command physically can do), approvals (decide what runs without asking), and checkpoints (undo what went wrong). Together they are the "input" layer (layer 1) and a big part of why delegation feels safe rather than terrifying.
Sandboxing: a fence around execution
A sandbox restricts what an executed command can touch (which files, whether it can reach the network) at the operating-system level, so even if the model asks for something destructive, the OS refuses. Codex uses the platform's native mechanism on each OS:
| Platform | Mechanism |
|---|---|
| macOS | sandbox-exec (Seatbelt) |
| Linux | Landlock |
| Windows | A purpose-built sandbox |
One subtle but important detail from the Codex write-up: the sandbox description applies only to Codex's own shell tool. Tools that come from MCP servers are not sandboxed by Codex; they are responsible for enforcing their own guardrails. That asymmetry is a real security consideration. When you add a third-party tool, you are trusting its author's safety practices, not your harness's sandbox. Codex also launched deliberately without general internet access, adding optional network access later under user control, on the principle that the smallest blast radius is the safest default.
Approvals: deciding what needs a human
You do not want to approve every ls. You also do not want a silent rm -rf or an unreviewed network call. So harnesses use a layered approval system. The "Inside the Agent Harness" piece describes Codex's tiers:
- Safe commands (
ls,cat,git status): auto-approved. - Pattern-matched commands: approved if they match a configured allowlist.
- Sandbox violations (network access, writes outside the workspace): require explicit approval.
- Granular policies: rules like "trust file writes but always prompt for network."
A module Codex calls the Guardian intercepts every tool call before execution, evaluates it against the policy, and either proceeds, prompts, or blocks. Claude Code exposes a similar idea as permission modes you cycle through:
| Mode | Behavior |
|---|---|
| Default | Ask before file edits and shell commands |
| Auto-accept edits | Edit files and run safe filesystem commands without asking; still prompt for the rest |
| Plan mode | Explore and propose a plan without editing source files |
| Auto mode | Evaluate all actions with background safety checks |
There is also project trust: whether project-local hooks and configs are even allowed to run, which protects you from a malicious repo you just cloned.
Smart approvals: removing the human bottleneck
A pure "ask the human every time" model breaks down for unattended runs (think CI, or a full-auto overnight task). The Codex deep-dive describes Smart Approvals: instead of interrupting a person, risky actions are routed through a guardian subagent that applies the policy rules and returns approve, deny, or escalate. The main agent keeps moving. Key properties: approvals can run in parallel, permissions persist across turns (granted once, honored for the session), and a spawned subagent inherits the parent's sandbox and network rules. This is the bridge between "safe but needs babysitting" and "safe enough to run on its own."
Checkpoints: undo for file changes
Even with sandboxing and approvals, the agent will sometimes make a change you do not want. Checkpoints are the undo button. Before Claude Code edits any file, it snapshots the current contents; press Escape twice (or ask) to rewind. Checkpoints are local to the session and separate from git. The important limitation: they only cover file changes. Actions that hit remote systems (databases, APIs, deployments) cannot be checkpointed, which is exactly why the harness asks before running commands with external side effects. You can undo a bad edit; you cannot un-send a deploy.