CHAPTER 14 · Cost, Latency, and Model Tiering · 5 / 13
Shrink and curate context
Since context is paid for on every loop iteration, controlling it is controlling cost:
- Reference, don't embed (Chapter 5). Don't paste documents into the prompt; let the model fetch via tools, paying for content only when used.
- Distill tool results. Cap long payloads, return decision-relevant fields, offer a "get more" tool. A giant unfiltered tool result is re-sent on every subsequent iteration, paying for it repeatedly.
- Curate history. Use a recent window or summarization rather than re-sending an ever-growing transcript.
- Keep prompts focused. Load task-specific instructions on demand instead of carrying every possible instruction in the system prompt on every call.