Skip to slide
Chapter 14 · Cost, Latency, and Model Tiering
135 / 191

CHAPTER 14 · Cost, Latency, and Model Tiering · 3 / 13

The master lever: model tiering

The highest-leverage optimization is routing each task to the cheapest model that can do it well (Chapter 4). A three-tier scheme:

  • Top tier for the primary interactive task where quality is paramount.
  • Mid tier for high-volume structured work (bulk extraction, classification) where you need throughput and lower cost and a capable-enough model suffices.
  • Low tier for trivial background tasks (titles, quick reformats, yes/no checks).

This often cuts cost dramatically and improves perceived speed, because the cheap models are also faster. Map every task to a tier explicitly; the default of "use the best model everywhere" is the costliest mistake.

← → arrow keys work too