Skip to slide
Chapter 14 · Cost, Latency, and Model Tiering
134 / 191

CHAPTER 14 · Cost, Latency, and Model Tiering · 2 / 13

Where latency comes from

Latency has overlapping sources:

  • Time to first token: how long before the model starts responding. Streaming hides this, but it still gates perceived responsiveness.
  • Generation time: proportional to output length and model speed.
  • Loop depth: each iteration is a serial round-trip; a deep loop is slow even if each call is fast.
  • Tool latency: slow external APIs (some take tens of seconds) stall the turn.
  • Context size: larger inputs take longer to process (prefill).
← → arrow keys work too