Skip to slide
Chapter 13 · Reliability: Retries, Idempotency, and Failure Handling
131 / 191

CHAPTER 13 · Reliability: Retries, Idempotency, and Failure Handling · 12 / 12

The reliability checklist

  • Parse all model output defensively; turn its mistakes into recoverable signals.
  • Guarantee one result per tool call.
  • Retry transient failures with capped exponential backoff and jitter; don't retry terminal errors.
  • Make side-effecting operations idempotent so retries are safe.
  • Persist progress so long work is resumable and input is never lost.
  • Classify each dependency as fatal or degradable, and handle accordingly.
  • Emit a clean error (not a hang) when a stream fails mid-flight.
  • Set timeouts on every external call; consider circuit breakers.
  • Cap iterations and detect stuck loops.
  • Log and measure failures so you can drive them down.

A reliable agent feels calm: when something underneath it breaks, the user gets a clear message or a degraded-but-useful result, finished work is preserved, and nothing hangs or duplicates. That calm is entirely the product of expecting failure and designing for it.

← → arrow keys work too