CHAPTER 15 · Observability and Evaluation
Observability and Evaluation
You cannot improve what you cannot see, and agents are unusually opaque: their behavior is probabilistic, multi-step, and dependent on context you assembled dynamically. This chapter covers two related disciplines, observability (seeing what an agent did) and evaluation (measuring whether it's any good). Together they turn "it seems to work" into "we know it works and we know when it regresses."