CHAPTER 08 · Parallelism, the Responses API, and the App Server · 2 / 8
The Responses API: built for agents
Codex migrated from the older Chat Completions API to the Responses API, and the reasons are instructive because they reveal what an agent actually needs from an API. The Codex deep-dive lists the wins:
- 40 to 80% better cache utilization, because the API was designed for the access patterns of agentic loops (the prefix-caching behavior from Chapter 5).
- A measurable SWE-bench improvement (around 3%), because better caching frees compute for reasoning within the same budget.
- Multi-turn tool use as a first-class shape, where Chat Completions was a GPT-3.5-era format not built for it.
- Parallel tool calls in a single response.
The API streams its response as Server-Sent Events (SSE): a sequence of small JSON events like response.output_text.delta (for streaming text to the UI) and response.output_item.added (objects the harness appends to the conversation for the next call). The harness consumes this stream and republishes it as internal events. This is why you see text appear token by token, and it is also how the harness detects that a tool call is needed mid-stream.
One more design choice worth knowing, because it explains a tradeoff: Codex deliberately does not use the API's previous_response_id feature, which would let the server remember the conversation so the client could send less. Codex keeps every request fully stateless (sending the whole history each time) to support Zero Data Retention customers, who cannot have their data stored server-side. The apparent inefficiency (re-sending everything) is bought back by prompt caching. It is a clean example of a real engineering tradeoff: statelessness and privacy in exchange for bigger requests, with caching making the exchange affordable.