CHAPTER 15 · Observability and Evaluation · 5 / 7
From observability to evaluation
Observability tells you what happened; evaluation tells you whether it was good. Evaluation is how you avoid the trap where a prompt tweak fixes one case and silently breaks five others.
Build an eval set
Curate a set of representative inputs with known-good expectations: realistic user messages (and documents) paired with what a correct response looks like. Cover the common cases, the tricky ones, and past bugs (regression cases). This set is your safety net; run it whenever you change a prompt, a tool, or a model.
What to measure
Agent quality is multi-dimensional. Useful metrics:
- Task success: did it accomplish the goal? (Often the hardest to measure automatically; may need a rubric.)
- Grounding / faithfulness: are factual claims actually supported by the source? For document agents, are citations correct (right location, verbatim quote)? This is checkable: verify each cited quote appears at the cited location.
- Format compliance: did it emit the required structured protocols correctly (parseable citation block, valid cell formats)? Deterministically checkable.
- Tool-use correctness: did it call the right tools, with valid arguments, in a sensible order? Did it avoid unnecessary calls?
- Refusal/safety behavior: does it refuse what it should and not over-refuse?
- Cost and latency: tokens and time per task (a quality regression can hide as a cost regression).
How to grade
- Deterministic checks where possible: does the citation block parse? Do cited quotes match the source? Did it call the expected tool? These are cheap and reliable; prefer them.
- LLM-as-judge for subjective quality: use a model to grade an answer against a rubric. Useful at scale, but calibrate it (judges have biases) and spot-check against human judgment.
- Human review for the highest-stakes or most subjective cases, and to validate your automated graders.
Run evals as part of change management
Treat prompts, tool definitions, and model choices as code: when you change them, run the eval set and compare to the baseline. A change that improves your target case but regresses others should be caught before it ships, not by users. This is the agent equivalent of a test suite, and it's what lets you iterate on a probabilistic system with confidence.