Skip to slide
Chapter 14 · Cost, Latency, and Model Tiering
145 / 191

CHAPTER 14 · Cost, Latency, and Model Tiering · 13 / 13

The cost/latency playbook

  1. Tier your models: cheapest capable model per task. (Biggest lever.)
  2. Reasoning off for bulk/one-shot work; on for hard interactive tasks.
  3. Shrink context: reference don't embed, distill tool results, curate history.
  4. Cut iterations: batch tools, good descriptions, cheap discovery tools.
  5. Parallelize independent work to cut wall-clock time.
  6. Cache external calls and exploit prompt caching with a stable prefix.
  7. Stream to mask the latency that remains.
  8. Right-size output and cap where bounded.
  9. Measure to find the real hotspots, then meter to bound per-user spend.

Cost and latency are won the same way: send fewer tokens to cheaper, faster models, fewer times, in parallel, and show progress while it happens.

← → arrow keys work too