Skip to slide
Chapter 14 · Cost, Latency, and Model Tiering
142 / 191

CHAPTER 14 · Cost, Latency, and Model Tiering · 10 / 13

Right-size output

Output tokens are expensive and slow. Don't ask for more than needed: cap output length for bounded tasks (a title needs a handful of tokens, not a paragraph), and instruct the model to be concise where verbosity adds no value. For structured extraction, the format constraints (Chapter 6) also keep output tight.

← → arrow keys work too