CHAPTER 18 · Reference Architecture and a Design Checklist · 11 / 15
Cost & performance
- Are tasks tiered to the cheapest capable model, with reasoning off for bulk? (Ch 4, 14)
- Do you reduce iterations (batching, good descriptions, discovery tools) and parallelize independent work? (Ch 14)
- Do you cache external calls and exploit prompt caching with a stable prefix? (Ch 9, 14)
- Is usage metered for per-user limits/attribution? (Ch 9, 14)