CHAPTER 04 · Model-Provider Abstraction and Multi-Model Strategy · 7 / 8
Reasoning/thinking as a per-call decision
Modern models can expose a reasoning/thinking mode. Treat it as a per-call toggle, not a global setting:
- On for interactive surfaces where the user benefits from seeing the agent think, or where the task is genuinely hard.
- Off for bulk and one-shot jobs where the reasoning stream just burns tokens and time.
Your neutral interface should carry an enableThinking-style flag, and each adapter translates it to that provider's mechanism (and, when off, explicitly zeroes the thinking budget where the provider allows, to actually save the tokens).