Skip to slide
Chapter 15 · Observability and Evaluation
149 / 191

CHAPTER 15 · Observability and Evaluation · 3 / 7

Tooling

Options range from rolling your own (structured logs keyed by a trace id, queryable in your log stack) to dedicated LLM-observability platforms that understand traces, spans, token costs, and prompt versions, and give you UIs to drill into a single turn or aggregate across many. Whatever you use, ensure:

  • A correlation id threads through the whole turn (and ideally the whole conversation) so you can reconstruct it end to end.
  • Token and cost capture per call, so you can attribute spend (Chapter 14).
  • Searchability by user, model, tool, outcome, so you can answer "which tool fails most?" or "what did this user's failing turn look like?"
← → arrow keys work too