Skip to slide
Chapter 14 · Cost, Latency, and Model Tiering
133 / 191

CHAPTER 14 · Cost, Latency, and Model Tiering · 1 / 13

Where cost comes from

LLM cost is driven by tokens: input (everything you send) and output (everything the model generates), usually priced separately with output more expensive. For an agent, the multipliers are:

  • Loop iterations. Each tool round-trip is another model call that re-sends accumulated context. A 5-iteration turn can cost several times a single call.
  • Context size. Every token of context is paid for on every call. Bloated prompts and dumped tool results are paid for repeatedly across the loop.
  • Model choice. Top-tier models can cost 10–50× a small model per token. Using a flagship model for a trivial task is the most common waste.
  • Reasoning tokens. Thinking/reasoning generates extra (billed) tokens. Valuable when it improves a hard answer; pure waste on a bulk job.
← → arrow keys work too