CHAPTER 14 · Cost, Latency, and Model Tiering · 10 / 13
Right-size output
Output tokens are expensive and slow. Don't ask for more than needed: cap output length for bounded tasks (a title needs a handful of tokens, not a paragraph), and instruct the model to be concise where verbosity adds no value. For structured extraction, the format constraints (Chapter 6) also keep output tight.