Skip to slide
Chapter 14 · Cost, Latency, and Model Tiering
139 / 191

CHAPTER 14 · Cost, Latency, and Model Tiering · 7 / 13

Parallelize independent work

When sub-tasks don't depend on each other, run them concurrently rather than serially. Bulk extraction across documents is embarrassingly parallel: fan out across documents instead of looping one at a time. This doesn't reduce total token cost, but it slashes wall-clock latency, which is often what users feel. (Mind provider rate limits when fanning out; Chapter 13.)

← → arrow keys work too