CHAPTER 16 · Production Realities: Economics, Security, and Governance · 1 / 10
The economics (and why they surprise people)
Recall the quadratic curve from Chapter 4: every turn replays the full history, and complex tasks fan out into many model calls. That shape drives a cost model that flat-rate-tool habits do not prepare you for. Codex moved to token-based billing in 2026: you pay for input tokens, plus cached input at roughly a tenth the rate, plus output tokens. Published 2026 analyses put a simple task near twelve cents, a complex one in the forty-to-sixty-five-cent range, and a debugging-heavy task higher still.
The danger is the loop. A flaky test or a circular dependency can send an agent into ten or twenty retries, each replaying the whole history, each more expensive than the last. The mitigation is exactly Chapter 12's budgets: cap turns and set token limits on automated runs so a retry storm cannot quietly run up a bill. The guide is blunt that cached input is the single most important cost lever, which means the Chapter 5 discipline (stable instructions, bounded scope, do not edit AGENTS.md mid-session) is not housekeeping; it is the cost model. A practical figure: roughly one hundred to two hundred dollars per developer per month at the team level.
There is also a striking efficiency contrast worth knowing: independent analyses report Claude Code tends to use more tokens per task and produce more thorough output, while Codex tends to be more concise; one widely cited build task reportedly used about 1.5 million tokens on Codex versus 6.2 million on Claude Code. Treat that as one data point, not a law, but it shows token efficiency varies a lot by tool and task.