CHAPTER 04 · Practical Guidance: When and How to Use Agentic RAG · 5 / 9
Lesson 5: Evaluation must account for process, not just outcomes
Most existing benchmarks judge only the final output and ignore how the system got there. For agentic systems, that is not enough. Meaningful evaluation needs process-level metrics: reasoning efficiency, tool usage patterns, and how well the system adapts to changing context. Without this, apparent improvements risk being anecdotal rather than systematic. This connects to the MAST work in the multi-agent topic, which is precisely a method for diagnosing process-level failures rather than just scoring outputs.