CHAPTER 10 · Document and File Processing Pipelines · 2 / 8
Format conversion is messy: isolate it
Converting between document formats (Office → PDF, etc.) is one of the messiest parts of any document system. Real-world files are malformed in creative ways: non-standard internal paths, corrupt structures, unusual encodings. Lessons:
- Use a battle-tested converter (a mature library or a tool like LibreOffice invoked as a subprocess) rather than writing your own. It encodes decades of edge-case handling.
- Normalise before converting. Pre-process known quirks (e.g. archives with non-standard internal path separators) so the converter doesn't choke. These fixes are unglamorous but essential once real user files arrive.
- Fail softly. If conversion fails, degrade: keep the document usable for text extraction even without a rendered preview, rather than rejecting the upload outright. Wrap conversion in a try/catch and continue with what succeeded.
- Treat the converter as an external dependency. It may need installing in your runtime image, it adds latency, and it can crash. Isolate it so its failures don't take down the request.