CHAPTER 14 · Cost, Latency, and Model Tiering · 3 / 13
The master lever: model tiering
The highest-leverage optimization is routing each task to the cheapest model that can do it well (Chapter 4). A three-tier scheme:
- Top tier for the primary interactive task where quality is paramount.
- Mid tier for high-volume structured work (bulk extraction, classification) where you need throughput and lower cost and a capable-enough model suffices.
- Low tier for trivial background tasks (titles, quick reformats, yes/no checks).
This often cuts cost dramatically and improves perceived speed, because the cheap models are also faster. Map every task to a tier explicitly; the default of "use the best model everywhere" is the costliest mistake.