Skip to slide
Chapter 14 · Cost, Latency, and Model Tiering
143 / 191

CHAPTER 14 · Cost, Latency, and Model Tiering · 11 / 13

Measure before optimizing

Don't guess where the cost goes; instrument it (Chapter 15). Track tokens per turn, per task type, and per model; track loop depth and tool latency. Usually a small number of patterns dominate the bill (a heavy tool result re-sent every iteration; a flagship model used for a background task). Find those and fix them, rather than micro-optimizing prompts that barely move the needle.

← → arrow keys work too