CHAPTER 08 · Parallelism, the Responses API, and the App Server · 8 / 8
Key takeaways
- Models can request several tool calls at once. Run independent read-only calls in parallel, but serialize mutating calls to avoid races; return results in order so
call_idlinkage stays intact. - The Responses API is built for agents: much better cache utilization, native multi-turn tool use, parallel calls, and SSE streaming.
- Codex stays stateless (re-sending full history) to support Zero Data Retention, and relies on prompt caching to make that affordable. A clean privacy-versus-efficiency tradeoff.
- The App Server exposes one harness to many surfaces over JSON-RPC, using three primitives: item, turn, thread, with a bidirectional flow that can pause for approvals.
- Reach for App-Server-style design when you go beyond a single CLI session (orchestration, IDEs, web).
Original sources: "Inside the Agent Harness," the "Inside the Codex Agent Loop" deep-dive, and OpenAI's Unlocking the Codex harness: how we built the App Server.