CHAPTER 14 · Cost, Latency, and Model Tiering · 2 / 13
Where latency comes from
Latency has overlapping sources:
- Time to first token: how long before the model starts responding. Streaming hides this, but it still gates perceived responsiveness.
- Generation time: proportional to output length and model speed.
- Loop depth: each iteration is a serial round-trip; a deep loop is slow even if each call is fast.
- Tool latency: slow external APIs (some take tens of seconds) stall the turn.
- Context size: larger inputs take longer to process (prefill).