Skip to slide
Chapter 15 · Observability and Evaluation
151 / 191

CHAPTER 15 · Observability and Evaluation · 5 / 7

From observability to evaluation

Observability tells you what happened; evaluation tells you whether it was good. Evaluation is how you avoid the trap where a prompt tweak fixes one case and silently breaks five others.

Build an eval set

Curate a set of representative inputs with known-good expectations: realistic user messages (and documents) paired with what a correct response looks like. Cover the common cases, the tricky ones, and past bugs (regression cases). This set is your safety net; run it whenever you change a prompt, a tool, or a model.

What to measure

Agent quality is multi-dimensional. Useful metrics:

  • Task success: did it accomplish the goal? (Often the hardest to measure automatically; may need a rubric.)
  • Grounding / faithfulness: are factual claims actually supported by the source? For document agents, are citations correct (right location, verbatim quote)? This is checkable: verify each cited quote appears at the cited location.
  • Format compliance: did it emit the required structured protocols correctly (parseable citation block, valid cell formats)? Deterministically checkable.
  • Tool-use correctness: did it call the right tools, with valid arguments, in a sensible order? Did it avoid unnecessary calls?
  • Refusal/safety behavior: does it refuse what it should and not over-refuse?
  • Cost and latency: tokens and time per task (a quality regression can hide as a cost regression).

How to grade

  • Deterministic checks where possible: does the citation block parse? Do cited quotes match the source? Did it call the expected tool? These are cheap and reliable; prefer them.
  • LLM-as-judge for subjective quality: use a model to grade an answer against a rubric. Useful at scale, but calibrate it (judges have biases) and spot-check against human judgment.
  • Human review for the highest-stakes or most subjective cases, and to validate your automated graders.

Run evals as part of change management

Treat prompts, tool definitions, and model choices as code: when you change them, run the eval set and compare to the baseline. A change that improves your target case but regresses others should be caught before it ships, not by users. This is the agent equivalent of a test suite, and it's what lets you iterate on a probabilistic system with confidence.

← → arrow keys work too