CHAPTER 13 · Reliability: Retries, Idempotency, and Failure Handling · 5 / 12
Persist progress for resumability
For multi-step or long-running work, persist each unit as it completes so a failure doesn't lose finished work:
- Save the user's input before starting (a failed turn doesn't lose their message).
- In a bulk job, write each result as it's produced, with a status field. A dropped connection or crash leaves completed units saved; a retry resumes from where it stopped by finding the still-
pendingunits.
This makes long jobs robust and is the same status-field discipline from Chapter 9, used here for recovery.