CHAPTER 15 · Observability and Evaluation · 3 / 7
Tooling
Options range from rolling your own (structured logs keyed by a trace id, queryable in your log stack) to dedicated LLM-observability platforms that understand traces, spans, token costs, and prompt versions, and give you UIs to drill into a single turn or aggregate across many. Whatever you use, ensure:
- A correlation id threads through the whole turn (and ideally the whole conversation) so you can reconstruct it end to end.
- Token and cost capture per call, so you can attribute spend (Chapter 14).
- Searchability by user, model, tool, outcome, so you can answer "which tool fails most?" or "what did this user's failing turn look like?"