CHAPTER 13 · Reliability: Retries, Idempotency, and Failure Handling · 12 / 12
The reliability checklist
- Parse all model output defensively; turn its mistakes into recoverable signals.
- Guarantee one result per tool call.
- Retry transient failures with capped exponential backoff and jitter; don't retry terminal errors.
- Make side-effecting operations idempotent so retries are safe.
- Persist progress so long work is resumable and input is never lost.
- Classify each dependency as fatal or degradable, and handle accordingly.
- Emit a clean error (not a hang) when a stream fails mid-flight.
- Set timeouts on every external call; consider circuit breakers.
- Cap iterations and detect stuck loops.
- Log and measure failures so you can drive them down.
A reliable agent feels calm: when something underneath it breaks, the user gets a clear message or a degraded-but-useful result, finished work is preserved, and nothing hangs or duplicates. That calm is entirely the product of expecting failure and designing for it.