CHAPTER 13 · Reliability: Retries, Idempotency, and Failure Handling · 11 / 12
Observability closes the loop
You can't improve reliability you can't see. Log failures with enough context (which turn, which tool, which document, the error) to diagnose them, and track failure rates so you know which dependency is the weak link (Chapter 15). Reliability work is iterative: instrument, find the top failure mode, fix it, repeat.