CHAPTER 10 · Document and File Processing Pipelines · 1 / 8
The ingestion pipeline
When a file is uploaded, a sequence of steps prepares it for use. A typical pipeline:
- Validate: accept only the formats you actually support; reject the rest immediately with a clear error. Narrowness is a feature: fewer formats means fewer edge cases.
- Create a record: write a metadata row with a
processingstatus before doing the heavy work, so you have a durable handle even if a later step fails. - Store the original bytes: put the raw file in object storage under a predictable, owner-scoped key. Keep the original untouched; everything else is derived.
- Produce a renderable form: convert to a viewable format (often PDF) if the source isn't already viewable, so your UI can display it consistently regardless of input type.
- Extract structure and text: pull a structural outline (headings/sections) and any metadata (page count) you'll need for navigation and citation.
- Record a first version: create the initial version row pointing at the stored bytes (Chapter 9).
- Mark ready: flip the status to
ready; onlyreadyitems are ever loaded into the agent's context.
If anything fails, mark the record error rather than leaving it half-built. The status field makes the pipeline's state explicit and recoverable.