CHAPTER 14 · Cost, Latency, and Model Tiering · 1 / 13
Where cost comes from
LLM cost is driven by tokens: input (everything you send) and output (everything the model generates), usually priced separately with output more expensive. For an agent, the multipliers are:
- Loop iterations. Each tool round-trip is another model call that re-sends accumulated context. A 5-iteration turn can cost several times a single call.
- Context size. Every token of context is paid for on every call. Bloated prompts and dumped tool results are paid for repeatedly across the loop.
- Model choice. Top-tier models can cost 10–50× a small model per token. Using a flagship model for a trivial task is the most common waste.
- Reasoning tokens. Thinking/reasoning generates extra (billed) tokens. Valuable when it improves a hard answer; pure waste on a bulk job.