CHAPTER 13 · Reliability: Retries, Idempotency, and Failure Handling · 3 / 12
Retries with backoff
For transient failures (timeouts, rate limits, 5xx from a provider or external API), retry, but correctly:
- Only retry idempotent or safe operations. A read is safe to retry; an action with side effects needs idempotency first (below).
- Exponential backoff with jitter. Wait progressively longer between attempts, with randomization, so you don't synchronize retries into a thundering herd against a recovering service.
- Cap attempts. A few retries, then surface a clear failure. Infinite retries just move the hang.
- Respect rate-limit signals. If a provider returns a retry-after, honor it rather than guessing.
Distinguish retryable errors (transient) from terminal ones (bad request, auth failure); retrying a terminal error just wastes time.