CHAPTER 15 · Observability and Evaluation · 6 / 7
Online signals
Beyond offline evals, watch production:
- Explicit feedback: thumbs up/down, corrections, regenerations. Aggregate by feature and model.
- Implicit signals: did the user accept the agent's proposed edit or reject it? Did they re-ask the same thing (a sign the first answer missed)? Did they abandon mid-turn?
- Failure and degradation rates: from your traces, which tools and dependencies fail most, and is it trending.
These tell you what your eval set can't: how the agent performs on the messy distribution of real use. Feed surprising production cases back into the eval set so it keeps reflecting reality.