CHAPTER 10 · Document and File Processing Pipelines · 6 / 8
The synchronous vs asynchronous decision
A central architectural choice: do you process the file during the upload request (synchronous) or hand it to a background worker (asynchronous)?
Synchronous (do it all in the request):
- Pros: simple, linear code; the response can return the finished, ready document; no queue infrastructure.
- Cons: the request is held open for the duration (conversion can be slow); large files risk timeouts; a burst of uploads ties up request handlers.
- Good when: files are modest in size, processing takes a few seconds, and volume is moderate.
Asynchronous (enqueue, process in a worker):
- Pros: uploads return instantly with a
processingstatus; heavy work runs off the request path; you can scale workers independently and retry failures. - Cons: more infrastructure (a queue, workers), and the client must poll or subscribe for completion; eventual-consistency UX.
- Good when: files are large, processing is slow (OCR, heavy conversion), or volume is high.
A pragmatic path: start synchronous, but design as if it will become asynchronous. The single thing that makes the later migration easy is already having the explicit status field (processing → ready → error). With that in place, moving steps 3–6 into a worker is a localised change; the rest of the system already understands "not ready yet." Don't build the queue before you need it; do build the status field from day one.