Skip to slide
Chapter 16 · Scaling and Infrastructure
155 / 191

CHAPTER 16 · Scaling and Infrastructure · 1 / 10

What's different about scaling an agent

Agents stress infrastructure in specific ways:

  • Long-lived connections. Streaming turns hold a connection open for many seconds. A server that assumes short requests will exhaust its connection pool.
  • Slow, external bottleneck. The dominant latency is the model provider, which you don't control. Your own compute is often nearly idle while waiting on the model.
  • Bursty, heavy work. File processing and bulk extraction spike CPU and memory unpredictably.
  • Rate limits upstream. Providers and external APIs cap your throughput; scaling your own fleet doesn't help if you hit their ceiling.

So scaling an agent is less about raw compute and more about concurrency, statelessness, and respecting upstream limits.

← → arrow keys work too