Skip to slide
Chapter 8 · Parallelism, the Responses API, and the App Server
56 / 142

CHAPTER 08 · Parallelism, the Responses API, and the App Server · 8 / 8

Key takeaways

  • Models can request several tool calls at once. Run independent read-only calls in parallel, but serialize mutating calls to avoid races; return results in order so call_id linkage stays intact.
  • The Responses API is built for agents: much better cache utilization, native multi-turn tool use, parallel calls, and SSE streaming.
  • Codex stays stateless (re-sending full history) to support Zero Data Retention, and relies on prompt caching to make that affordable. A clean privacy-versus-efficiency tradeoff.
  • The App Server exposes one harness to many surfaces over JSON-RPC, using three primitives: item, turn, thread, with a bidirectional flow that can pause for approvals.
  • Reach for App-Server-style design when you go beyond a single CLI session (orchestration, IDEs, web).

Original sources: "Inside the Agent Harness," the "Inside the Codex Agent Loop" deep-dive, and OpenAI's Unlocking the Codex harness: how we built the App Server.

← → arrow keys work too