What is retrieval-augmented generation (RAG)?
Retrieval-augmented generation (RAG) is a way of building AI systems. In it, relevant documents are first retrieved from an outside source and passed to a language model. The model uses them to write its answer instead of relying only on what it learned in training.
Also known as: RAG, retrieval augmented generation
Researched and fact-checked by AI, with no human review. 7 sources listed below. How we verify
Last updated
How it works
The name comes from a 2020 paper by Patrick Lewis and 11 co-authors. Their models joined a pre-trained text generator to a searchable index of Wikipedia. The authors reported state-of-the-art results on three open-domain question-answering tasks.
The glossary of a March 2025 report from the U.S. National Institute of Standards and Technology (NIST) describes RAG more generally. It says a generative AI model is coupled with a separate retrieval system, also called a knowledge base. That system finds material relevant to each user query and supplies it to the model as context.
Preparing a knowledge base commonly involves splitting long documents into chunks and converting them into embeddings. Embeddings are numerical representations that allow search by meaning. Microsoft's documentation for its Azure AI Search product describes both steps.
Why it matters
A September 2020 post from Facebook AI, now Meta AI, said a RAG system's knowledge can be changed by swapping the documents it retrieves from, without retraining the model.
Retrieval does not remove errors. Stanford researchers tested RAG-based legal research tools in a preprint study. They reported in May 2024 that Lexis+ AI and Ask Practical Law AI gave incorrect information more than 17% of the time. They put Westlaw's AI-Assisted Research at more than 34%, though the tools made fewer errors than the general-purpose GPT-4.
Retrieval also adds a security exposure. The NIST report says a RAG knowledge base can be poisoned so that the model gives an attacker's chosen output for particular queries.
Where things stand in 2026
Stanford's 2026 AI Index lists RAG alongside function calling and text embedding as key capabilities for deployed applications. It says standard pipelines, which fetch individual text chunks, can struggle with questions that require combining information across documents.
Google presents larger context windows as an alternative for some tasks. Its long-context guide, last updated in June 2026, says small windows led to RAG's rapid adoption. It says Gemini models can instead be given all relevant material up front. But it adds that retrieval remains valuable in specific scenarios. The AI Index cautions that bigger windows do not guarantee better results.
Microsoft's documentation, last updated in August 2026, contrasts classic single-query RAG with what it calls agentic retrieval. In that approach, a language model splits a complex question into subqueries that run in parallel. Microsoft recommends agentic retrieval for new projects on its service.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, arXiv (Lewis et al.; accepted at NeurIPS 2020)
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, National Institute of Standards and Technology (NIST AI 100-2e2025)
- Retrieval Augmented Generation: Streamlining the creation of intelligent natural language processing models, Meta AI (published as Facebook AI)
- AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries, Stanford Institute for Human-Centered AI (HAI)
- AI Index Report 2026, Chapter 2: Technical Performance, Stanford Institute for Human-Centered AI (HAI)
- Retrieval-augmented generation (RAG) in Azure AI Search, Microsoft Learn
- Long context, Google AI for Developers (Gemini API documentation)