CHAPTER 16 · Scaling and Infrastructure · 5 / 10
Respect and manage upstream rate limits
You can scale your fleet infinitely and still be capped by the model provider's rate limits. Manage them deliberately:
- Concurrency control / queuing toward providers so you don't blow past limits and trigger errors.
- Backoff and retry on rate-limit responses (Chapter 13).
- Spread load across providers or accounts where appropriate; tiering (Chapter 14) also helps by sending bulk work to higher-throughput smaller models.
- Cache to avoid redundant calls (Chapter 9).
When fanning out parallel work (bulk extraction), bound the fan-out to stay within limits rather than firing thousands of simultaneous calls.