Writing

What Is RAG, and Why Can't You Just Give an AI All Your Documents?

· AI · Engineering Notes, Local AI, RAG

Drafted June 2025. Finished and published August 2026 as part of the migration away from WordPress.

While experimenting with local AI, RAG kept coming up. I understood the basic claim: it lets a model answer questions about your own documents. I did not understand why it needed a separate retrieval system.

My question was much simpler: why not put every document into the prompt and ask the model to find the answer?

The context window is a budget

A model can only consider a limited number of tokens in one request. That limit is its context window.

If the documents fit comfortably, giving the model everything can be the simplest option. The problem is that a collection grows. More input takes more memory and processing time, and eventually some of it has to be cut. Even before that point, the useful sentence may be buried among thousands of irrelevant ones.

As of June 2025, the OpenAI and Anthropic API models relevant to this question look roughly like this:

Provider and modelsContext windowApproximate dense A4 pages
OpenAI GPT-4o and GPT-4o mini128,000 tokens190
OpenAI o3, o3-pro and o4-mini200,000 tokens300
OpenAI GPT-4.1 family1,000,000 tokens1,500
Anthropic Claude 3.7 Sonnet and Claude 4200,000 tokens300

The page figures use OpenAI’s estimate of 100 tokens for about 75 English words and assume 500 words on a dense, single-spaced A4 page. They are only there to show scale. Font size and spacing can change them considerably.

GPT-4.1’s one-million-token window was the outlier. Most of the other models in the table sat between 128,000 and 200,000 tokens, or roughly 190 to 300 dense pages. That is a lot for one prompt, but it is still a limit rather than a document library. The instructions, question and answer also need room, and sending the whole collection again for every question adds processing time and cost.

RAG stands for retrieval-augmented generation (Lewis et al., 2020). Instead of making the model search the whole collection inside its prompt, it adds a search step first:

  1. split documents into smaller passages;
  2. find the passages most relevant to the question; and
  3. give only those passages to the language model.

The model still writes the answer. RAG just chooses the evidence it gets to read.

I tried two kinds of search. Vector search turns passages and questions into lists of numbers called embeddings, then looks for passages with similar meaning. BM25 is keyword search, which is useful when the query contains an exact name or code. I also combined their rankings into a hybrid result.

A small test

I generated 120 synthetic policy documents in code. Twelve contained invented facts that the model could not already know. The other 108 contained similar departments, dates, colours and identifiers to make retrieval less obvious.

There were twelve questions: six paraphrased the source and six asked for an exact name or code. One looked like this:

Question: What does incident code ORCHID-17 mean?
Source:   Incident code ORCHID-17 denotes a cold-storage checksum mismatch.

I used Qwen3-8B as the generator, Qwen3-Embedding-0.6B for vectors, Qdrant for local storage and BM25 for keyword search. The generator had a native 32,768-token context window. I capped its input at 30,000 tokens to leave room for the answer, equal to roughly 45 dense A4 pages under the estimate above.

Each question ran under five conditions: no documents, as much of the corpus as would fit in a 30,000-token prompt, the top five vector matches, the top five BM25 matches, and the top five hybrid matches.

The full-corpus test ran in three different document orders. The other conditions ran once.

MethodCorrect answersGold source in top fiveMedian input tokensMedian elapsed time
No documents0 / 12N/A20.53.26 s
Stuff as much of the corpus as fits13 / 36N/A30,0004.17 s
Vector search, top five11 / 1211 / 121,833.50.48 s
BM25, top five11 / 1212 / 121,8290.48 s
Hybrid search, top five11 / 1212 / 121,8300.48 s

With no documents, the model got none of the invented answers right.

The giant prompt was surprisingly unreliable. Its three document orders scored 6, 3 and 4 out of 12. Seven attempts lost their source at the 30,000-token cut, but truncation was not the whole explanation. The correct source survived in 29 attempts, and the model still answered only 13 of them correctly.

The retrieval methods used about 1,830 input tokens instead of 30,000. All three scored 11 out of 12 and finished in a median of 0.48 seconds rather than 4.17 seconds.

For the ORCHID-17 question, vector, BM25 and hybrid search all selected the correct document. The model returned:

Incident code ORCHID-17 denotes a cold-storage checksum mismatch.

What I had missed

I had thought of RAG as a way to give a model more information. It is really a way to be selective about information.

The test also showed that retrieval and answering are separate problems. Dense search missed the correct source for one question. BM25 and hybrid search found it, but the generator still returned NOT FOUND. A document appearing in the top five did not guarantee that the model would use it.

There was no clear winner between the three search methods here. BM25 handled exact identifiers and also found the one semantic source that vector search missed. All three produced the same final score.

My original idea was not completely wrong. If a collection is small and fits easily, sending it all can avoid the extra machinery. Anthropic’s 2024 contextual retrieval write-up makes a similar point for knowledge bases below roughly 200,000 tokens (Anthropic).

Once the collection becomes large, changes often or needs source tracking, retrieval starts to make sense. It lets the model spend its context window on likely evidence rather than every document available.

This was a synthetic corpus with twelve questions, not a general benchmark. It was enough to answer my early question. You can give an AI all your documents when they fit. RAG is what you build when choosing the right documents first becomes more useful than sending everything.

← All posts