CHAPTER 04 · Model-Provider Abstraction and Multi-Model Strategy · 6 / 8
Model tiering: the multi-model strategy
Once you can call any model trivially, route work to the right model. A common and effective scheme is three tiers:
- Top tier: the most capable (and expensive) models, for the primary interactive task where quality matters most: complex reasoning, drafting, the main chat.
- Mid tier: fast, capable-enough models for high-volume structured work: bulk extraction, classification, where you run many calls and need throughput and lower cost.
- Low tier: the cheapest, fastest models for trivial background tasks: generating a title, a quick reformat, a yes/no check.
Map tasks to tiers deliberately. Interactive chat → top tier with reasoning on. Bulk extraction → mid tier with reasoning off. Title generation → low tier. This single discipline can cut costs dramatically while improving perceived performance, because cheap tasks finish faster.
A nice refinement: when users bring their own keys, route background tasks to whichever provider they actually have a key for, picking that provider's cheapest model. The tiering becomes "cheapest available model of an available provider."