Skip to slide
Chapter 14 · Cost, Latency, and Model Tiering
137 / 191

CHAPTER 14 · Cost, Latency, and Model Tiering · 5 / 13

Shrink and curate context

Since context is paid for on every loop iteration, controlling it is controlling cost:

  • Reference, don't embed (Chapter 5). Don't paste documents into the prompt; let the model fetch via tools, paying for content only when used.
  • Distill tool results. Cap long payloads, return decision-relevant fields, offer a "get more" tool. A giant unfiltered tool result is re-sent on every subsequent iteration, paying for it repeatedly.
  • Curate history. Use a recent window or summarization rather than re-sending an ever-growing transcript.
  • Keep prompts focused. Load task-specific instructions on demand instead of carrying every possible instruction in the system prompt on every call.
← → arrow keys work too