What is RAG?
Last reviewed 7 October 2026
Retrieval-augmented generation (RAG) is a method in which a system first retrieves relevant passages from a collection of documents and then gives them to a language model as context for its answer. The model writes from the retrieved text rather than only from what it absorbed in training. The name comes from a 2020 paper by Lewis and colleagues, which paired a pre-trained sequence-to-sequence model with a dense vector index searched by a neural retriever.
Also called: retrieval-augmented generation, retrieval augmented generation, RAG pipeline, grounded generation.
Why Retrieval-augmented generation (RAG) matters
A language model knows only what was in its training data, and that data has a cut-off date and almost never includes your contracts, your runbooks or last week's pricing decision. Retraining a model each time a policy changes is slow and costly. RAG separates the two jobs: the model keeps its general language ability, and a retrieval step supplies the facts at the moment of the question.
It also makes answers checkable. When the system keeps track of which passages it retrieved, it can show them next to the answer, and a reader can open the source and judge for themselves. AWS describes RAG in these terms: the model references an authoritative knowledge base outside its training data before it responds.
How it works
Indexing happens before any question is asked. Documents are split into small chunks, an embedding model turns each chunk into a vector, and the vectors go into an index together with the chunk text and metadata such as the source file, the date and who may read it. Many systems also keep a keyword index, because exact terms such as product codes match poorly on vectors alone.
At question time the query is embedded in the same way, and the index returns the chunks whose vectors sit closest to it, often merged with keyword results and re-ranked. The system then builds a prompt that contains the question, the selected chunks and an instruction to answer only from them and to cite them. The model generates the answer from that prompt.
The original paper described two variants: RAG-Sequence, which uses the same retrieved passages for the whole answer, and RAG-Token, which can draw on different passages for each generated token. Most production systems today are simpler: they retrieve once, place the text in the prompt and call a general-purpose model. Either way, the index has to be refreshed as documents change, or the answers go stale with it.
Worked example: a leave policy question
An employee asks the HR assistant how many days of carer's leave they can take. The retriever searches the indexed HR handbook and returns three chunks: the carer's leave section, a paragraph on how leave interacts with sick pay, and an older version of the same policy that was never removed from the folder.
The prompt asks the model to answer only from the chunks and to name the section it used. The model answers from the current policy and cites it. The stale version is the real lesson: retrieval returns whatever is in the index, so the index needs an owner who removes superseded documents, and the access list on each chunk has to match the access list on the source file, or the assistant will quote a document the employee was never allowed to open.
How Swfte relates to it
Built in the product
Cortex, the Swfte desktop app, includes knowledge bases with retrieval that run on the device. You can add files, folders, URLs and notes, and with embeddings from a local model through Ollama the indexing and the answer can both stay on the machine. Cortex can also answer with local models through Ollama or LM Studio, or with a hosted model you choose.
The company brain is a separate, customer-hosted design. Its graph of people, groups and accounts is built. Permission-aware search over documents, with hybrid keyword and vector retrieval filtered by each caller's access, exists as libraries but its endpoints are in progress, so it is not something to rely on yet. For a step-by-step build that does not depend on Swfte at all, read the RAG guide.
Related terms
- Agentic RAG
Agentic RAG is retrieval-augmented generation in which an AI agent controls the retrieval step, deciding what to search for, which sources to use and whether to search again before it answers.
- Knowledge graph
A knowledge graph is a store of facts in which entities, such as people, products or systems, are nodes and the named relationships between them are edges.
- Semantic layer
A semantic layer is software that sits between stored data and the tools that query it, translating tables and columns into named business terms such as revenue, active customer or churn.
- Model Context Protocol (MCP)
The Model Context Protocol (MCP) is an open protocol that gives AI applications a standard way to connect to external tools, data sources and prompts, so one integration can work with any compatible client.
- AI agent
An AI agent is a software program that uses an AI model to decide what to do next and then acts through tools to reach a goal it was given.
Common questions
- What does RAG stand for in AI?
- RAG stands for retrieval-augmented generation. Retrieval means searching a collection of documents for passages that match the question. Augmented means those passages are added to the model's prompt. Generation is the model writing its answer from that enlarged prompt rather than from training data alone.
- Does RAG stop a model from making things up?
- It reduces the problem but does not remove it. If retrieval returns the wrong passage, or nothing useful, the model can still produce a confident answer. Good systems instruct the model to say when the sources do not cover the question, show citations, and test retrieval quality separately from answer quality.
- Do I need a vector database for RAG?
- Not always. Small collections can be searched with keyword search or an in-memory index. A vector index helps when wording varies between the question and the document. Many teams combine vector and keyword search, because exact identifiers such as invoice numbers often match better on keywords.
- When should I not use RAG?
- Skip it when the answer does not depend on your documents, such as rewriting text or general reasoning. It also fits badly when you need the model to learn a style or format consistently, which is a case for fine-tuning or better prompts, and when the source data changes faster than you can index it.
Sources
Definitions on this page were read on the sources below on 7 October 2026. Where sources define the term differently, the page says so. The full glossary lists more terms.
- Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, arXiv 2005.11401 (read 2026-10-07)
- AWS, What is retrieval-augmented generation (read 2026-10-07)