CHAPTER 14 · Cost, Latency, and Model Tiering · 13 / 13
The cost/latency playbook
- Tier your models: cheapest capable model per task. (Biggest lever.)
- Reasoning off for bulk/one-shot work; on for hard interactive tasks.
- Shrink context: reference don't embed, distill tool results, curate history.
- Cut iterations: batch tools, good descriptions, cheap discovery tools.
- Parallelize independent work to cut wall-clock time.
- Cache external calls and exploit prompt caching with a stable prefix.
- Stream to mask the latency that remains.
- Right-size output and cap where bounded.
- Measure to find the real hotspots, then meter to bound per-user spend.
Cost and latency are won the same way: send fewer tokens to cheaper, faster models, fewer times, in parallel, and show progress while it happens.