CHAPTER 16 · Production Realities: Economics, Security, and Governance · 2 / 10
Where agents excel and where they struggle
Honesty here separates a useful guide from a sales pitch. The "Everything About Codex" guide maps it cleanly.
| Agents are strong at | Agents struggle with |
|---|---|
| Refactoring at scale (mechanical, test-verified) | Architecture decisions (cross-system trade-offs) |
| Documentation (code exists, summarize it) | Ambiguous requirements (it picks an interpretation rather than asking) |
| Testing (generate and iterate until green) | Deep domain knowledge (rules not in the code) |
| Bug fixing for well-scoped, reproducible defects | Security-sensitive code (subtle, plausible-looking flaws) |
| Repo onboarding and PR creation | Cross-system dependencies (complexity in the seams) |
The pattern: agents win where work is verifiable and tedious and struggle where it needs judgment, ambiguity, or context outside the repo. The failure modes follow: they can hallucinate APIs that look right, produce silent errors that pass a weak test suite, and because the output is fluent, they invite over-trust. The "Deep Research of Deep Research" survey gives this its name: jagged intelligence, brilliant at some hard tasks, surprisingly bad at simpler adjacent ones.