What is RAG? Retrieval-augmented generation, explained without the jargon

Ask a general AI model about your company's refund policy and it will give you a confident, well written answer. It will also probably be wrong, because it has never seen your refund policy. Retrieval-augmented generation, or RAG, is the standard fix: before the model answers, look up the relevant documents and hand them over.

That is the whole idea. The rest of this guide explains how it works under the bonnet, when RAG is the right tool and when it is not, why it still matters now that models can read a million tokens at once, and how to build one, with or without code.

RAG in one paragraph

Think of an exam. A model without RAG is sitting it closed-book: it can only use what it memorised during training, and if it does not know, it may bluff. A model with RAG is sitting it open-book. A search step finds the most relevant pages from your own material and puts them in front of the model along with the question, and the model writes its answer from those pages. The name describes the three steps exactly: retrieve the relevant text, augment the prompt with it, then generate the answer.

Why models need it

A large language model learns from a huge snapshot of text and then stops learning. That leaves three gaps that RAG fills.

RAG also reduces made-up answers, because the model is told to answer from the text in front of it and to say so when that text does not contain the answer. It does not eliminate them, which is covered below.

How RAG works, step by step

A RAG system has two halves: a preparation stage you run whenever your documents change, and a question stage that runs every time someone asks something.

Preparation: building the index

  1. Collect the documents. PDFs, web pages, wiki articles, tickets, spreadsheets, whatever the answers live in.
  2. Split them into chunks. A 60-page manual is too much to hand over whole for every question, so it is cut into passages of a few hundred words, usually with a little overlap so ideas are not sliced in half.
  3. Turn each chunk into an embedding. An embedding model converts text into a long list of numbers that captures its meaning. Passages about similar things end up with similar numbers, even if they use different words. "How do I get my money back" and "refund policy" land close together.
  4. Store them in a vector database. The embeddings, plus the original text and details such as the source and date, go into a database built to find "nearest" matches quickly.

Question time: retrieve, augment, generate

  1. Embed the question with the same embedding model.
  2. Retrieve the handful of chunks whose meaning is closest to the question, typically somewhere between three and twenty.
  3. Augment the prompt: the model receives an instruction ("Answer using only the sources below, and cite them"), the retrieved chunks and the question.
  4. Generate the answer, ideally with references back to the chunks it used.

In code, the question stage is short. This sketch shows the shape of it, whichever libraries you use:

question = "What is our refund window for annual plans?"

# 1. Retrieve: find the passages closest in meaning to the question
hits = vector_db.search(embed(question), top_k=5)

# 2. Augment: put them in the prompt with clear instructions
sources = "\n\n".join(f"[{h.id}] {h.text}" for h in hits)
prompt = f"""Answer using only the sources below. Cite source ids.
If the answer is not in the sources, say you don't know.

{sources}

Question: {question}"""

# 3. Generate
answer = llm.generate(prompt)

The short version: RAG is search plus a writer. The search finds the right pages, the model reads them and writes the answer. If the search finds the wrong pages, the best model in the world will still give a bad answer.

RAG vs fine-tuning vs long context

RAG is one of several ways to get a model to use knowledge it was not trained on. They solve different problems, and mixing them up is the most common and most expensive mistake teams make.

ApproachWhat it changesBest forWeak at
RAGWhat the model can see for each questionLarge or changing knowledge, answers with sourcesQuestions that need the whole picture at once
Fine-tuningHow the model behavesA consistent style, format or specialist taskTeaching facts, which go stale and get muddled
Long contextHow much you paste inOne-off analysis of a few big documentsLarge collections, cost and speed at scale
Tools and MCPWhat the model can look up or do liveLive data such as calendars, CRMs and databasesSearching unstructured documents by meaning

The rule of thumb: fine-tuning teaches a model how to behave, RAG gives it something to read. If the problem is that the model does not know your facts, fine-tuning is the wrong tool. Our guide on how to fine-tune an LLM covers the cases where it is the right one.

Doesn't a million-token context window make RAG pointless?

Several 2026 models, including DeepSeek V4, Kimi K3 and Google's Gemini line, can read around a million tokens in one go, roughly a few thousand pages. For a handful of long documents you only need once, pasting them in is often simpler and better than building a RAG system.

It does not replace RAG for everything else. Most organisations have far more than a million tokens of material. Sending everything with every question is slow and, even with caching, much more expensive than sending five relevant passages. Models also still pay less attention to material buried in the middle of a huge prompt. In practice, the two now work together: retrieval picks the most relevant material, and a large context window means you can afford to send more of it.

Where MCP fits

The Model Context Protocol lets a model call tools, including search tools. A RAG search can be one of those tools, so the model decides when to look something up rather than searching on every question. MCP is the plumbing that connects the model to a source. RAG is the technique for finding the right passages inside that source.

Where RAG goes wrong

Most disappointing RAG systems fail at the search step, not the model. The usual culprits:

What good RAG looks like in 2026

The basic version above is a starting point. The systems that work well in production usually add some of the following.

How to build one

Without writing code

You may already have a RAG system and not know it. Uploading files to a Claude Project, a ChatGPT project or custom GPT, or Google's NotebookLM, and asking questions about them, uses retrieval behind the scenes. For a team knowledge base of a few hundred documents this is often enough, and it is the fastest way to find out whether the idea works for you before anyone builds anything. Many workplace tools, from help desks to intranets, now include an AI search option built the same way. If you are weighing up which of these to pay for, our guide to AI tools for small business can help.

With code

A first working version takes a developer a day or two. The usual ingredients:

Start with one well-defined collection, such as your help centre, and a list of real questions people ask. Get those answered well, with sources, before you add anything else. Getting a good result out of the final step is still partly a prompting job, so the techniques in our prompt engineering basics guide apply here too.

Frequently asked questions

What does RAG stand for in AI?

RAG stands for retrieval-augmented generation. It describes a system that retrieves relevant information from a set of documents, adds (augments) it to the prompt, and has a language model generate an answer from it.

What is RAG in simple terms?

RAG is an open-book exam for AI. Instead of answering from memory, the model is first given the most relevant passages from your own documents and writes its answer from those, ideally citing which passage each part came from.

What is the difference between RAG and fine-tuning?

Fine-tuning changes how a model behaves by training it further on examples, which suits a consistent style or specialist task. RAG leaves the model unchanged and gives it relevant documents to read for each question, which suits large or frequently changing knowledge. For teaching a model your facts, RAG is almost always the better choice.

Is RAG still needed with long-context models?

Yes, for most real uses. Million-token context windows are excellent for analysing a few long documents, but most organisations have far more material than that, and sending everything with every question is slow and expensive. Retrieval picks the relevant passages and long context lets you send more of them.

Does RAG stop AI hallucinations?

It reduces them but does not stop them. A model answering from retrieved text is far more accurate than one answering from memory, but it can still misread a passage or fill gaps. Requiring citations, instructing the model to say when the sources do not contain the answer, and testing against real questions all help.

What is a vector database?

A vector database stores embeddings, which are numerical representations of the meaning of text, and finds the ones closest to a query very quickly. It is the search engine at the heart of most RAG systems. Popular options include pgvector for Postgres, Pinecone, Weaviate, Qdrant and Chroma.

What to take from this

RAG is the most practical way to make an AI useful on your own information, and it is less exotic than the acronym suggests: good search, then a model that reads the results. Almost all the quality comes from the search half, so spend your effort on clean documents, sensible chunking, hybrid search and a set of test questions. Try it first with the file upload in a tool you already pay for. If the answers are useful, you have a business case. If they are not, you have saved yourself a project.

Learn what the acronyms actually mean.

Every weekday morning before 8am we explain what changed in AI and whether it matters for your work. Plain English, no hype, free.

You're in. First issue lands tomorrow morning.