Skip to slide
Chapter 14 · Cost, Latency, and Model Tiering
140 / 191

CHAPTER 14 · Cost, Latency, and Model Tiering · 8 / 13

Cache aggressively

  • Cache external API calls (Chapter 9): repeated identical lookups become instant and free, and you stay under rate limits.
  • Exploit provider prompt caching where available: keeping a stable prefix (system prompt, tool definitions) constant lets the provider cache it and charge less for repeated input. Structure your context so the invariant parts come first.
← → arrow keys work too