CHAPTER 16 · Scaling and Infrastructure · 1 / 10
What's different about scaling an agent
Agents stress infrastructure in specific ways:
- Long-lived connections. Streaming turns hold a connection open for many seconds. A server that assumes short requests will exhaust its connection pool.
- Slow, external bottleneck. The dominant latency is the model provider, which you don't control. Your own compute is often nearly idle while waiting on the model.
- Bursty, heavy work. File processing and bulk extraction spike CPU and memory unpredictably.
- Rate limits upstream. Providers and external APIs cap your throughput; scaling your own fleet doesn't help if you hit their ceiling.
So scaling an agent is less about raw compute and more about concurrency, statelessness, and respecting upstream limits.