Skip to slide
Chapter 7 · Retrieval: RAG vs Tools vs Long Context
60 / 191

CHAPTER 07 · Retrieval: RAG vs Tools vs Long Context · 1 / 6

The three strategies

1. Long context ("just include it")

Put the relevant documents directly into the prompt and let the large context window hold them. Simple, no infrastructure.

  • Wins when: the total content is small relative to the window, the task needs the whole document (not a snippet), and you can afford the tokens.
  • Loses when: content is large (cost and latency balloon), there's a lot of it (the model gets "lost in the middle" and quality drops), or it changes often (you re-send it every call).

2. Retrieval-augmented generation (RAG)

Pre-process documents into chunks, embed them into a vector store, and at query time retrieve the most semantically similar chunks to inject into context.

  • Wins when: you have a large corpus, queries are answerable from fragments, and you need to search across many documents by meaning rather than exact text.
  • Loses when: the task needs whole-document understanding (chunking fragments the structure), exactness matters (embeddings retrieve "similar," not "the exact clause"), or you need verbatim quotes with precise locations (chunk boundaries and approximate retrieval undermine citation precision). RAG also adds real infrastructure: an embedding pipeline, a vector DB, chunking strategy, and retrieval tuning.

3. Tool-based retrieval ("let the model fetch")

Give the model tools to read and search content on demand, "read this document," "find this phrase," "list what's available," and let it decide what to pull, driven by the task.

  • Wins when: the model can reason about what it needs (it knows it wants "the termination clause"), exactness matters, you need whole-document or targeted reads, and content changes (tools always fetch the current version).
  • Loses when: the corpus is so large the model can't even know what exists without help (then you add a search/list tool, blending toward RAG), or when latency from multiple fetch round-trips is unacceptable.
← → arrow keys work too