CHAPTER 07 · Retrieval: RAG vs Tools vs Long Context · 2 / 6
Why tools are often the right default for document agents
For an agent working with a bounded, identified set of documents (a user's files, a project folder, a matter), tool-based retrieval tends to win for several reasons:
- Exactness and citations. When the model reads the actual document text (paginated, complete) rather than approximate chunks, it can quote verbatim and cite precise locations. For domains where citation accuracy is the whole point (law, medicine, finance), this is decisive.
- Freshness. Tools fetch the current version every time, so edits and updates are reflected automatically. A vector index, by contrast, must be re-embedded when content changes.
- Grounding. Forcing the model to explicitly fetch content makes it engage with the real text instead of leaning on training-data priors.
- Simplicity. No embedding pipeline, no vector store, no chunking strategy to tune. The model's own reasoning is the retrieval policy.
- Whole-document tasks. Summarising or editing a contract needs the whole thing in coherent order, which chunked retrieval handles poorly.
The cost is multiple round-trips (the model reads, then acts), which the agent loop already handles, and which you mitigate with batching tools ("fetch these N documents at once") and cheap targeted tools ("find this phrase" instead of reading everything).