Retrieval for the enterprise
The RAG pipeline: ten stages, the decision at each, and what to log
A stage-by-stage view of a retrieval-augmented generation pipeline, with the decision to make, the failure to expect and the evidence to keep at each stage.
A RAG pipeline fetches passages from your own data at question time and gives them to a model with the question. It has ten stages: ingest, parse, chunk, embed, index, retrieve, rerank, assemble context, generate and evaluate. Each stage has a decision and a way to fail quietly. In an enterprise the stage that matters most is retrieval, because it must return only what the asking person may read.
Last verified 2026-10-07. Sources are listed at the end of the page.
What is a RAG pipeline?
The term comes from a 2020 paper by Lewis and colleagues. It describes models that combine “parametric and non-parametric memory”: a pre-trained language model that stores knowledge in its weights, and an external index it can look things up in. In that paper the external index was a dense vector index of Wikipedia, read through a neural retriever.
Framework documentation splits the work into two phases. LangChain’s guide describes indexing (load, split, embed, store) and then retrieval and generation, where relevant splits are fetched and a model answers from a prompt that holds both the question and the retrieved data. LlamaIndex lists loading, indexing, storing, querying and evaluation. This page uses ten stages so that each decision has its own row.
What are the stages of a RAG pipeline, and what do you decide at each?
Read the stages in order. Most quality problems start early and show up late.
1. Ingest
Decide which sources feed the index and how often each is re-read. Copy the source system’s access list with every document, as well as its text. Record the source, the version and the time of reading.
2. Parse
Decide how tables, scans, slides and attachments become text. A table flattened into a run of numbers is a common silent loss. Keep the original file reference so an answer can point back to it.
3. Chunk
Decide the unit of retrieval: a paragraph, a section, a page. Small chunks match precisely but lose context. Anthropic’s contextual retrieval post addresses this by prepending chunk-specific explanatory context to each chunk before embedding.
4. Embed
Decide the embedding model and where it runs. Documents and queries must use the same model. Changing the model means re-embedding the whole corpus, so record the model name with each vector.
5. Index
Decide the store and the index type. The pgvector documentation says an HNSW index has better query performance than IVFFlat for the speed and recall trade-off, but slower builds and more memory. Without an index, pgvector performs exact search with perfect recall.
6. Retrieve
Decide the search mode and the filter. Keyword search finds exact terms, vector search finds meaning, and hybrid search combines them. The pgvector README points to Postgres full-text search with Reciprocal Rank Fusion for that. Apply the access filter here, before ranking.
7. Rerank
Decide whether a second model re-scores a wider candidate set. In Anthropic’s test the pipeline retrieved 150 candidates and reranked them to the top 20. Reranking adds latency and a model call, so measure what it buys on your questions.
8. Assemble context
Decide how many chunks to send, in what order, with what labels. Mark retrieved text as data. Research on long contexts found that performance often degrades when relevant information sits in the middle of a long input, so order and count matter.
9. Generate
Decide the prompt, the model and the citation rule. Require the answer to cite the chunks it used and to say when the retrieved text does not contain the answer.
10. Evaluate
Decide what you score. Score retrieval separately from the answer: did the right chunk arrive in the top results, and did the answer stay inside it? Keep a fixed question set and re-run it after every change to a stage.
What can go wrong at each stage, and what should you log?
A pipeline you cannot inspect cannot be fixed. This table gives the minimum record per stage.
| Stage | What can go wrong | What to log |
|---|---|---|
| Ingest | Stale index: a source changed or was deleted and the index still holds the old text. Access list not copied, or copied once and never refreshed. | Source, document version, read time, access list as read, and a count of documents added, changed and removed per run. |
| Parse | Tables, headers or scanned pages lost or garbled. | Parser name and version, pages or elements dropped, and a sample of parsed text per source type. |
| Chunk | A chunk loses the context that gives it meaning, such as which product or year it refers to. | Chunking settings, chunk count per document and the chunk identifier used in answers. |
| Embed | Query and document vectors from different models. A model change not followed by a re-embed. | Embedding model name and version stored with each vector, plus failed batches. |
| Index | Approximate search misses relevant rows. Filtering applied after the index scan returns too few results. | Index type and parameters, index build time and the recall you measured on a test set. |
| Retrieve | A person receives chunks they may not read. A relevant chunk is never returned. | Caller identity, the filter applied, chunk identifiers returned and their scores. Never the chunk text if the text is restricted. |
| Rerank | The reranker drops the one right chunk, or adds latency for no gain. | Candidates in, candidates out and the score change per chunk. |
| Assemble context | Retrieved text carries instructions that the model follows. The context is too long and the key chunk is buried. | Final chunk list in order, token count and any chunk flagged by an injection check. |
| Generate | The answer states something the chunks do not support, or cites a chunk it did not use. | Model and version, prompt version, cited chunk identifiers and the answer. |
| Evaluate | A test set that no longer resembles real questions, so scores look good while users are unhappy. | Test-set version, scores per stage and the date of each run. |
The logging column follows OWASP’s recommendation (LLM08:2025) to keep detailed, immutable logs of retrieval activity so suspicious behaviour can be found and answered.
How does retrieval respect the source system’s permissions?
A RAG index is a second copy of your documents. Unless it also carries who may read each one, it becomes a way round the access controls of the system the document came from. OWASP’s entry on vector and embedding weaknesses names this directly: inadequate or misaligned access controls can lead to unauthorised access to embeddings containing sensitive information. Its mitigation is permission-aware vector stores with strict logical partitioning of the data.
Three rules follow. First, copy the source access list at ingest and refresh it on a schedule, because people change teams and groups change membership. Second, filter by the caller’s identity inside the query, before results are ranked or returned, not in application code afterwards. Third, treat a missing or incomplete access list as “nobody may read this”, never as “everybody may”.
The second rule also affects recall. The pgvector README warns that with approximate indexes, filtering is applied after the index is scanned: if a condition matches 10% of rows, with HNSW and the default ef_search of 40, only about 4 rows will match on average. It recommends iterative index scans to fetch more. A permission filter is a condition like any other, so test it on users who can read very little.
What are the failure modes that hurt most in production?
- Stale index. The model answers confidently from a policy that was replaced last month. Fix with change-driven ingest, deletion handling and a freshness date shown beside each citation.
- Access-list leak. A person sees text from a document they cannot open in the source system. Fix with filtering inside the query, tests with low-privilege users and logging of every filter applied.
- Injection through retrieved text. LangChain’s documentation states that retrieved documentation may contain text that resembles instructions, and that no prompt or delimiter strategy fully prevents indirect prompt injection. Limit what the model can do after reading, and validate outputs.
- Poisoned source. OWASP notes that poisoning can be intentional or unintentional. Accept data only from trusted sources and review what enters the index.
- Silent retrieval miss. The right chunk exists and is never returned. Only a retrieval-level test set finds this.
What did Anthropic measure with contextual retrieval?
Anthropic’s post, published on 19 September 2024, reports that contextual embeddings reduced the top-20-chunk retrieval failure rate by 35%, that combining contextual embeddings with contextual BM25 reduced it by 49%, and that adding reranking reduced it by 67%. These are Anthropic’s results on its own test datasets. The post recommends testing configurations for your use case, and this page does not claim the same gains for your data.
The post also gives a one-time cost of $1.02 per million document tokens to generate the contextualised chunks, as published on that page when read on 2026-10-07. Treat the figure as a pointer to the vendor’s page, not as a price for your model or your corpus.
Where Swfte fits: what is built and what is not
Status words used on this site: Built, In progress, Roadmap. No dates are given.
| Capability | Status | Note |
|---|---|---|
| Directory sync and the evidence-backed graph of people, groups and reporting lines | Built | The structured part of the company brain. It is not document retrieval. |
| Per-object access lists copied from the source, with filtering inside SQL | In progress | Libraries with tests exist. An unknown or incomplete access list makes the object invisible. The HTTP endpoints are not finished. |
| Chunker, embedding client and hybrid keyword plus vector search with reciprocal-rank fusion | In progress | Same state: libraries exist, endpoints are not finished, so document search and ask your company are not available yet. |
| On-device RAG knowledge bases in Cortex | Built | A desktop app that indexes files, folders and URLs on the device, with a local path through Ollama embeddings. |
| Cortex reading the company brain | Roadmap | Designed for. Not available. |
If you have one corpus, one permission model and a small team, you do not need Swfte for this: a framework and Postgres with pgvector will do. Our step-by-step guide is how to build a RAG system. The brain pages explain why the permission model is the hard part. See access and governance.
Sources and last verified
Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.
- Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (arXiv 2005.11401). Original definition of RAG and its parametric and non-parametric memory.
- Anthropic, Introducing Contextual Retrieval. Contextual embeddings, contextual BM25, reranking, and the published failure-rate reductions and cost.
- LlamaIndex, high-level concepts. The stages of a RAG application as the framework lists them.
- LangChain, build a RAG agent. Indexing and retrieval phases, and the warning on indirect prompt injection.
- pgvector README. Exact versus approximate search, HNSW versus IVFFlat, filtering after the index scan, and hybrid search with Reciprocal Rank Fusion.
- OWASP, LLM08:2025 Vector and Embedding Weaknesses. Access control, poisoning and logging guidance for vector stores.
- Liu et al., Lost in the Middle (arXiv 2307.03172). Position of relevant information in long contexts and its effect on performance.
Frequently asked questions
What is a RAG pipeline?
A RAG pipeline is the sequence of steps that finds relevant passages in your own data and gives them to a language model with the question. It covers ingest, parse, chunk, embed, index, retrieve, rerank, context assembly, generation and evaluation. The original 2020 paper paired a pre-trained model with a vector index of Wikipedia.
Which stage of a RAG pipeline matters most?
Retrieval, because the model can only answer from what it is given. A wrong or forbidden chunk at that stage cannot be repaired later. Chunking and the access filter come next, since they decide what retrieval can find and who may see it. Measure retrieval on its own, apart from the generated answer.
Do I need a vector database for RAG?
Not always. The pgvector documentation shows Postgres can hold vectors, run exact search by default and add approximate indexes, and it can be combined with Postgres full-text search for hybrid retrieval. A dedicated vector database makes sense when scale or features exceed what your existing database serves. Test with your own data first.
How do I stop RAG from leaking restricted documents?
Copy each document’s access list into the index at ingest, refresh it on a schedule and filter by the caller’s identity inside the retrieval query. Treat a missing access list as no access. Then test with accounts that can read very little, and log the filter applied to every query.
Is hybrid search better than vector search alone?
Often for exact terms such as part numbers, names and codes, where keyword matching helps. Anthropic’s test combined contextual embeddings with contextual BM25 and reported a larger drop in retrieval failures than embeddings alone, on its own datasets. Whether that holds for your data is an empirical question, so compare on a fixed question set.
Is the Swfte company brain a RAG system today?
Not yet. The brain’s graph of people, groups and reporting lines is built. Document search with per-object access lists and hybrid retrieval is in progress: libraries exist and the endpoints are not finished. Cortex has on-device RAG knowledge bases today, but it does not read the brain yet.