Skip to slide
Chapter 15 · Observability and Evaluation
152 / 191

CHAPTER 15 · Observability and Evaluation · 6 / 7

Online signals

Beyond offline evals, watch production:

  • Explicit feedback: thumbs up/down, corrections, regenerations. Aggregate by feature and model.
  • Implicit signals: did the user accept the agent's proposed edit or reject it? Did they re-ask the same thing (a sign the first answer missed)? Did they abandon mid-turn?
  • Failure and degradation rates: from your traces, which tools and dependencies fail most, and is it trending.

These tell you what your eval set can't: how the agent performs on the messy distribution of real use. Feed surprising production cases back into the eval set so it keeps reflecting reality.

← → arrow keys work too