# How to build a RAG system

Canonical: https://www.swfte.com/how-to-build-a-rag-system
Last verified: 2026-10-06
Difficulty: Intermediate
Time: About 90 minutes to a working system, plus time to collect your own documents
Cost: No software cost. Everything runs on your machine; the only spend is electricity and disk space.
Hardware: A laptop with 8 GB of memory is enough for the 3B chat model used here. More memory lets you use a larger chat model; retrieval itself is light.

## Short answer

A RAG system finds the passages in your documents that match a question and gives only those to a language model, so it can answer and cite them. Build it as six small parts: chunk, embed, store, retrieve, answer, test. This guide uses Postgres with pgvector, Ollama and about 150 lines of Python, and puts the permission check inside the search query, not after it.

## Who this is for

- Developers who want a RAG system they understand part by part before choosing a framework.
- Teams whose documents are private and who need answers to respect who is asking.
- Anyone who has seen a RAG demo give a confident wrong answer and wants a way to measure it.

Not for:
- You need answers that join facts across many systems or reason over relationships between people, services and decisions. Start with a [company brain](https://www.swfte.com/how-to-build-a-company-brain-for-ai) design.
- You want to teach a model a new style or format. That is fine-tuning, not retrieval.

## Prerequisites

- Docker installed and running, and Python 3.10 or later (the psycopg docs list 3.10 to 3.15).
- Ollama installed from ollama.com/download. About 2 GB free for the chat model and under 1 GB for the embedding model.
- A folder of Markdown files you are allowed to index. The guide creates three small sample files so you can test permissions safely.
- Comfort reading Python and basic SQL. Nothing here needs a GPU.

## Before you start: is RAG the right tool?

Use RAG when the answer lives in your documents, the documents change, and you need to show where an answer came from. Retrieval keeps the knowledge outside the model, so you can update a file and the next answer changes. It also lets you delete a document and be sure it is gone from future answers.

RAG is the wrong tool when the problem is behaviour rather than knowledge: a fixed output format, a tone of voice, a narrow classification. Fine-tuning handles those better. The [fine-tuning guide](https://www.swfte.com/how-to-fine-tune-an-llm-on-your-own-data) covers that choice, and the architecture options are laid out in the [RAG architecture article](https://www.swfte.com/blog/rag-llm-architecture-implementation-guide-2026).

**The parts of a RAG system and the choice made in this guide**

| Part | What it does | Choice here |
| --- | --- | --- |
| Chunker | Cuts documents into passages a model can use | Paragraph packing, 1,200 characters |
| Embedder | Turns text into vectors for search by meaning | nomic-embed-text through Ollama |
| Store | Holds text, vectors and permissions together | PostgreSQL with pgvector |
| Retriever | Finds the best passages for a question | Keyword and vector search, fused by rank |
| Generator | Writes the answer from the passages | llama3.2 through Ollama |
| Test set | Tells you whether a change helped | Hand-written questions with expected sources |

## Steps

### Step 1: Start Postgres with pgvector

Outcome: A Postgres 18 database with the pgvector extension available on localhost:5432.

pgvector adds a vector column type and similarity search to Postgres. Using it means your text, vectors and permission data live in one database, so one SQL query can search by meaning, by keyword and by who is asking. That is the reason this guide uses it instead of a separate vector database.

The command below starts the image the pgvector project documents. Change the password; `mysecretpassword` is the placeholder from the Postgres image documentation. The volume path is `/var/lib/postgresql` because the Postgres image documentation says PostgreSQL 18 and later use that path, not `/var/lib/postgresql/data`.

On a shared machine, publish the port on localhost only rather than on every interface (see Docker's `--publish` documentation).

Start the database:

```bash
docker run --name rag-db -e POSTGRES_PASSWORD=mysecretpassword -p 5432:5432 -v rag-pgdata:/var/lib/postgresql -d pgvector/pgvector:pg18-trixie
```

> NOTE: The pgvector README lists images for Postgres 13 to 18. If you already run another Postgres version, pick the matching tag from the README instead.

### Step 2: Create a Python environment and pull the models

Outcome: An isolated environment with the three libraries installed and two models downloaded.

Create a virtual environment so the libraries stay out of your system Python. Install psycopg (the Postgres driver), pgvector (the Python helpers that register the vector type) and ollama (the client for the local model server).

Then pull the two models. `nomic-embed-text` is an embedding-only model of 274 MB per its Ollama library page. `llama3.2` is the 3B chat model, listed at 2.0 GB. Any chat model that Ollama serves works; set the environment variable `RAG_CHAT_MODEL` later to try another.

Environment and libraries:

```bash
python3 -m venv .venv
source .venv/bin/activate
pip install "psycopg[binary]" pgvector ollama
```

Models:

```bash
ollama pull nomic-embed-text
ollama pull llama3.2
```

> TIP: Ollama listens on http://localhost:11434 once it is running. If `ollama pull` cannot connect, start the Ollama app or run `ollama serve`.

### Step 3: Create three sample documents with different readers

Outcome: A docs folder where each sub-folder name is the group allowed to read what is inside it.

The sample set has one document for everyone, one for engineering and one for HR only. The folder name becomes the permission group in the next step. This is a stand-in for a real source system, where the access list comes from the file store or the directory, and you copy it at ingest time.

Keep the sample small so you can read every result. Later, point the ingest script at your own folder of Markdown files and keep the one-folder-per-group layout, or replace the group rule with a lookup against your own access data.

Create the folders:

```bash
mkdir -p docs/everyone docs/engineering docs/hr
```

docs/everyone/leave-policy.md:

```text
# Leave policy

Every employee receives 25 days of annual leave per calendar year, plus public holidays.

Leave requests go to your line manager at least two weeks ahead, except for sickness.
```

docs/engineering/deploy-runbook.md:

```text
# Deploy runbook

Production deploys run from the main branch after two approvals.

If a deploy fails the health check, roll back with the previous release tag and open an incident.
```

docs/hr/salary-bands.md:

```text
# Salary bands

A senior engineer is in band E4. Band E4 pays between 70,000 and 90,000 euros per year.

These bands are confidential to HR.
```

### Step 4: Chunk, embed and store the documents

Outcome: Every document is cut into passages, turned into vectors and stored with its text and its allowed groups.

Save the shared helper first. It opens the connection, creates the extension if needed and registers the vector type, in the same order the pgvector examples use. It also holds the two embedding functions. The `search_document:` and `search_query:` prefixes are the task prefixes the pgvector Ollama example adds for nomic-embed-text; skip them and retrieval quality drops.

Then save the ingest script. The chunker is deliberately simple: it splits on blank lines and packs paragraphs up to 1,200 characters. A paragraph longer than that stays whole, which is acceptable for a first version and is the first thing to change if your documents contain long unbroken text. The script copies the group onto every chunk, creates a full-text index, an index on the groups and an HNSW index for cosine distance, and deletes a document's old chunks before inserting new ones, so running it again after an edit replaces rather than duplicates.

The vector dimension is 768, matching nomic-embed-text in the pgvector example. If you swap the embedding model, check the length of one vector and change `DIMENSIONS`.

rag_common.py:

```python
import os

import ollama
import psycopg
from pgvector import Vector
from pgvector.psycopg import register_vector

DSN = os.environ.get('RAG_DSN', 'postgresql://postgres:mysecretpassword@localhost:5432/postgres')
EMBED_MODEL = 'nomic-embed-text'
CHAT_MODEL = os.environ.get('RAG_CHAT_MODEL', 'llama3.2')
DIMENSIONS = 768  # nomic-embed-text produces 768-dimensional vectors


def connect():
    conn = psycopg.connect(DSN, autocommit=True)
    conn.execute('CREATE EXTENSION IF NOT EXISTS vector')
    register_vector(conn)
    return conn


def embed_documents(texts):
    # nomic-embed-text expects a task prefix on documents and on queries
    result = ollama.embed(model=EMBED_MODEL, input=['search_document: ' + t for t in texts])
    return result.embeddings


def embed_query(text):
    return ollama.embed(model=EMBED_MODEL, input='search_query: ' + text).embeddings[0]
```

ingest.py:

```python
import re
import sys
from pathlib import Path

from rag_common import DIMENSIONS, Vector, connect, embed_documents


def chunk(text, max_chars=1200):
    """Split on blank lines, then pack paragraphs into chunks of up to max_chars."""
    chunks, current = [], ''
    for block in re.split(r'\n\s*\n', text):
        block = block.strip()
        if not block:
            continue
        if current and len(current) + len(block) + 2 > max_chars:
            chunks.append(current)
            current = ''
        current = current + '\n\n' + block if current else block
    if current:
        chunks.append(current)
    return chunks


def main(folder):
    conn = connect()
    conn.execute(f"""
        CREATE TABLE IF NOT EXISTS chunks (
            id bigserial PRIMARY KEY,
            doc_id text NOT NULL,
            source text NOT NULL,
            position int NOT NULL,
            content text NOT NULL,
            allowed_groups text[] NOT NULL,
            ingested_at timestamptz NOT NULL DEFAULT now(),
            embedding vector({DIMENSIONS})
        )""")
    conn.execute("CREATE INDEX IF NOT EXISTS chunks_fts ON chunks USING GIN (to_tsvector('english', content))")
    conn.execute('CREATE INDEX IF NOT EXISTS chunks_acl ON chunks USING GIN (allowed_groups)')
    conn.execute('CREATE INDEX IF NOT EXISTS chunks_vec ON chunks USING hnsw (embedding vector_cosine_ops)')

    files = sorted(Path(folder).rglob('*.md'))
    total = 0
    for path in files:
        group = path.parent.name  # the folder name stands in for "who may read this"
        doc_id = str(path)
        chunks = chunk(path.read_text(encoding='utf-8'))
        conn.execute('DELETE FROM chunks WHERE doc_id = %s', (doc_id,))  # re-ingest replaces, never duplicates
        for start in range(0, len(chunks), 32):
            batch = chunks[start:start + 32]
            for offset, (content, emb) in enumerate(zip(batch, embed_documents(batch))):
                conn.execute(
                    'INSERT INTO chunks (doc_id, source, position, content, allowed_groups, embedding) '
                    'VALUES (%s, %s, %s, %s, %s, %s)',
                    (doc_id, path.name, start + offset, content, [group], Vector(emb)),
                )
        total += len(chunks)
    print(f'ingested {total} chunks from {len(files)} files')


if __name__ == '__main__':
    main(sys.argv[1])
```

Run it:

```bash
python ingest.py docs
```

The script prints one line (each sample file fits in one chunk):

```text
ingested 3 chunks from 3 files
```

### Step 5: Retrieve with keyword and vector search together

Outcome: A search function that returns the best passages a given group may read.

Vector search finds passages that mean the same thing as the question. Keyword search finds exact terms such as a part number or a policy name that embeddings blur. Running both and fusing the ranks is called hybrid search, and the pgvector README names reciprocal rank fusion (RRF) as a way to combine them. RRF adds 1 divided by (60 + rank) from each list, so a passage ranked high in both lists wins without you having to compare two different score scales.

The SQL is adapted from the pgvector-python hybrid example. Two changes matter. The keyword side uses `websearch_to_tsquery`, which the Postgres documentation says never raises a syntax error on raw user input. And both sub-queries carry `allowed_groups && %(groups)s`, the array overlap operator, so a chunk the caller may not read is never a candidate. The `hnsw.iterative_scan` setting makes the vector index keep scanning when the filter removes rows; without it a filtered query can return fewer rows than you asked for.

Try it with different groups. An engineer should not see the salary document, however well it matches.

retrieve.py:

```python
from rag_common import Vector, connect, embed_query

HYBRID_SQL = """
WITH semantic AS (
    SELECT id, RANK() OVER (ORDER BY embedding <=> %(emb)s) AS rank
    FROM chunks
    WHERE allowed_groups && %(groups)s::text[]
    ORDER BY embedding <=> %(emb)s
    LIMIT 20
),
keyword AS (
    SELECT id, RANK() OVER (ORDER BY ts_rank_cd(to_tsvector('english', content), q) DESC) AS rank
    FROM chunks, websearch_to_tsquery('english', %(q)s) q
    WHERE to_tsvector('english', content) @@ q
      AND allowed_groups && %(groups)s::text[]
    ORDER BY ts_rank_cd(to_tsvector('english', content), q) DESC
    LIMIT 20
)
SELECT c.id, c.source, c.content,
       COALESCE(1.0 / (%(k)s + semantic.rank), 0.0) + COALESCE(1.0 / (%(k)s + keyword.rank), 0.0) AS score
FROM semantic
FULL OUTER JOIN keyword ON semantic.id = keyword.id
JOIN chunks c ON c.id = COALESCE(semantic.id, keyword.id)
ORDER BY score DESC
LIMIT %(n)s
"""


def search(conn, query, groups, n=5, k=60):
    """Hybrid search: vector rank and keyword rank fused with reciprocal rank fusion.
    The group filter sits inside both sub-queries, so a chunk the caller may not
    read is never a candidate."""
    conn.execute('SET hnsw.iterative_scan = strict_order')
    params = {'emb': Vector(embed_query(query)), 'q': query, 'groups': groups, 'k': k, 'n': n}
    rows = conn.execute(HYBRID_SQL, params).fetchall()
    return [{'id': r[0], 'source': r[1], 'content': r[2], 'score': float(r[3])} for r in rows]


def rerank(query, hits, keep=3):
    """Optional second pass with a cross-encoder."""
    from sentence_transformers import CrossEncoder

    model = CrossEncoder('cross-encoder/ms-marco-MiniLM-L6-v2')
    ranks = model.rank(query, [h['content'] for h in hits])
    ranks = sorted(ranks, key=lambda r: r['score'], reverse=True)
    return [hits[r['corpus_id']] for r in ranks[:keep]]


if __name__ == '__main__':
    import sys

    conn = connect()
    groups = sys.argv[2].split(',')
    for hit in search(conn, sys.argv[1], groups):
        print(f"{hit['score']:.4f}  {hit['source']}")
```

Same question, two different readers:

```bash
python retrieve.py "what is the salary band for a senior engineer" everyone,engineering
python retrieve.py "what is the salary band for a senior engineer" everyone,hr
```

**What you should see**

| Caller groups | Salary document returned? | Why |
| --- | --- | --- |
| everyone,engineering | No | The chunk's allowed_groups is [hr], which does not overlap |
| everyone,hr | Yes, ranked first | Keyword and meaning both match, and the group overlaps |

### Step 6: Generate an answer that cites its sources

Outcome: A command that answers a question using only retrieved passages and names them.

Number the retrieved passages, tell the model to use only those, ask for citations in square brackets, and give it an exact sentence to use when the sources do not contain the answer. That last instruction matters: without a permitted way to say "not found", a model fills the gap from its training data and the answer looks cited when it is not.

The script uses `ollama.generate`, as the pgvector Ollama example does, and prints the sources it passed in. Always show users the sources. It is the cheapest way to let them catch a wrong answer.

ask.py:

```python
import sys

import ollama

from rag_common import CHAT_MODEL, connect
from retrieve import rerank, search

PROMPT = """You answer questions using only the numbered sources below.
Cite the sources you used as [1], [2]. If the sources do not contain the answer,
reply exactly: I could not find that in the documents you are allowed to see.

Sources:
{sources}

Question: {question}
Answer:"""


def ask(question, groups, use_rerank=False):
    conn = connect()
    hits = search(conn, question, groups, n=6)
    if use_rerank and hits:
        hits = rerank(question, hits, keep=3)
    else:
        hits = hits[:3]
    sources = '\n\n'.join(f"[{i}] ({h['source']}) {h['content']}" for i, h in enumerate(hits, start=1))
    answer = ollama.generate(model=CHAT_MODEL, prompt=PROMPT.format(sources=sources, question=question)).response
    return answer, hits


if __name__ == '__main__':
    answer, hits = ask(sys.argv[1], sys.argv[2].split(','))
    print(answer)
    print('\nSources used:', ', '.join(h['source'] for h in hits))
```

Ask as an engineer:

```bash
python ask.py "How many days of annual leave do I get?" everyone,engineering
python ask.py "What is the salary band for a senior engineer?" everyone,engineering
```

### Step 7: Optionally add a reranker

Outcome: A second pass that reorders the top candidates with a model that reads question and passage together.

A bi-encoder embedding scores a question and a passage separately. A cross-encoder reads them together and is usually better at ordering a short list, at the cost of running a model per candidate. Retrieve more than you need (the script asks for six), rerank, then keep three.

Install sentence-transformers (its documentation recommends Python 3.10 or later and PyTorch 2.2 or later) and run ask.py with reranking switched on from Python. The model `cross-encoder/ms-marco-MiniLM-L6-v2` has 22.7 million parameters and an Apache 2.0 licence on its model card. Add reranking only after the next step shows retrieval, not generation, is your weak point; it adds latency and a dependency.

Install:

```bash
pip install -U sentence-transformers
```

Try it from a Python shell:

```python
from ask import ask
answer, hits = ask('How do I roll back a failed deploy?', ['everyone', 'engineering'], use_rerank=True)
print(answer)
```

### Step 8: Write a test set and measure retrieval

Outcome: A repeatable score for retrieval quality and a hard check that permissions hold.

Write 20 to 50 real questions with the file that should answer each. Save them as `eval.jsonl`, one JSON object per line. Include at least a few "forbid" lines: questions whose best answer sits in a document the caller must not see. Those are your permission tests, and they fail loudly if a filter is dropped.

The script reports hit@3 (is the expected file in the top three?), mean reciprocal rank (how high is it, on average?) and the number of permission leaks, and exits non-zero on any leak so you can run it in continuous integration. Measure before you change anything, change one thing, measure again. Chunk size, the number of candidates and the RRF constant are all worth one experiment each.

Retrieval scores do not tell you whether the answer is faithful. Read 20 generated answers against their cited passages by hand, and mark each as supported, partly supported or unsupported. Do it again after each change to the prompt or the model.

eval.jsonl:

```json
{"question": "How many days of annual leave do I get?", "groups": ["everyone", "engineering"], "expect": "leave-policy.md"}
{"question": "How do I roll back a failed deploy?", "groups": ["everyone", "engineering"], "expect": "deploy-runbook.md"}
{"question": "What is the salary band for a senior engineer?", "groups": ["everyone", "engineering"], "forbid": "salary-bands.md"}
{"question": "What is the salary band for a senior engineer?", "groups": ["everyone", "hr"], "expect": "salary-bands.md"}
```

eval.py:

```python
import json
import sys

from rag_common import connect
from retrieve import search

# eval.jsonl: one JSON object per line, for example
# {"question": "How many days of annual leave do I get?", "groups": ["everyone", "engineering"], "expect": "leave-policy.md"}
# {"question": "What is the salary band for a senior engineer?", "groups": ["everyone", "engineering"], "forbid": "salary-bands.md"}


def main(path):
    conn = connect()
    hits_at_3, reciprocal, scored, leaks = 0, 0.0, 0, 0
    for line in open(path, encoding='utf-8'):
        case = json.loads(line)
        results = search(conn, case['question'], case['groups'], n=5)
        sources = [r['source'] for r in results]
        if 'expect' in case:
            scored += 1
            if case['expect'] in sources[:3]:
                hits_at_3 += 1
            if case['expect'] in sources:
                reciprocal += 1.0 / (sources.index(case['expect']) + 1)
            else:
                print('MISS', case['question'], sources)
        if 'forbid' in case and case['forbid'] in sources:
            leaks += 1
            print('LEAK', case['question'], sources)
    print(f'hit@3 {hits_at_3}/{scored}   MRR {reciprocal / max(scored, 1):.2f}   permission leaks {leaks}')
    sys.exit(1 if leaks else 0)


if __name__ == '__main__':
    main(sys.argv[1])
```

Run it:

```bash
python eval.py eval.jsonl
```

The last line has this shape (the numbers depend on your documents):

```text
hit@3 3/3   MRR 1.00   permission leaks 0
```

### Step 9: Keep the index fresh and the permissions current

Outcome: A plan for re-ingesting changed files and for changes to who may read what.

Re-run `ingest.py` when files change. Because it deletes a document's old chunks before inserting new ones, it is safe to run on a schedule. Deleting a source file should delete its chunks too: add a pass that removes rows whose `doc_id` no longer exists on disk, and test it.

Permissions are the part that rots. If someone leaves a group, the chunk still carries the old list until the next ingest. Re-sync access lists on a shorter schedule than content, or look them up at query time instead of copying them, and record the time of the last sync next to each chunk. When a permission could not be read, treat the document as invisible, not public.

For anything beyond one team, the same design extends to a [company brain](https://www.swfte.com/how-to-build-a-company-brain-for-ai): more sources, identity resolution, and evidence on every answer.

## When to stop tuning and change approach

If retrieval hits the right passage in your top three results for most questions and the answers are still wrong, the problem is the prompt or the model, not the search. If the right passage is missing from the top ten, fix chunking or add keyword search before touching the model. Do not move to a bigger model to compensate for a retrieval fault: it costs more and hides the cause.

If questions need several documents combined, or answers depend on who owns what, plain retrieval will keep failing in the same way. That is the point to read about [agentic RAG](https://www.swfte.com/agentic-rag) or a structured knowledge layer, and to price it with the [cost of RAG](https://www.swfte.com/cost-of-rag) breakdown. If you are weighing retrieval against tool access through MCP, [RAG vs MCP](https://www.swfte.com/compare/rag-vs-mcp) explains the difference.

## Troubleshooting

| Symptom | Likely cause | Fix |
| --- | --- | --- |
| psycopg.OperationalError: connection refused (or "could not connect to server") | The container is not running, or the port is not published. | Run `docker ps` and check `rag-db` is listed with 5432 published. If it exited, run `docker logs rag-db`. |
| The extension "vector" is not available | You started a plain Postgres image, not the pgvector one. | Remove the container and start `pgvector/pgvector:pg18-trixie` as in step 1. Use the same command, with a fresh volume name if the old one holds a database created by another image. |
| ConnectionError when calling ollama.embed | The Ollama server is not running or is not on localhost:11434. | Start the Ollama app or run `ollama serve`, then retry. Confirm the models exist with `ollama list`. |
| expected 768 dimensions, not N | The embedding model changed after the table was created. | Drop the table, set `DIMENSIONS` to the length of one vector from the new model, and ingest again. Never mix vectors from two models in one column. |
| A filtered query returns fewer than 5 rows although more chunks match | The approximate index scanned its candidate list, then the group filter removed rows. | Keep `SET hnsw.iterative_scan = strict_order` as in retrieve.py. The pgvector README describes iterative scans for exactly this case. For very selective filters, a partial index per group or partitioning is the README's other suggestion. |
| The answer is fluent but not in the sources | The prompt allows the model to fall back on what it already knows, or retrieval returned the wrong passages. | Print the sources. If they are wrong, work on retrieval. If they are right, tighten the prompt and run a smaller test of "not found" questions that must return the exact refusal sentence. |
| Exact terms such as an invoice number are never found | Embeddings blur identifiers, and the keyword side tokenises them differently from your text. | Check that the keyword query can match the term on its own: run the `websearch_to_tsquery` condition in psql. If the term is split by the English dictionary, store a second text column with a simple configuration and index that. |

## Verify it worked

- [ ] ingest.py prints a chunk count and running it twice leaves the same count, not double.
- [ ] retrieve.py returns the salary document for the hr group and never for the engineering group.
- [ ] ask.py answers the leave question with a citation like [1] and lists leave-policy.md under sources.
- [ ] ask.py returns the exact "could not find" sentence for a question none of the allowed documents answer.
- [ ] eval.py prints permission leaks 0 and exits with status 0; deleting the group filter from the SQL makes it fail (try it once, then restore it).
- [ ] Editing a sample file and re-running ingest.py changes the next answer.

## Next steps

- [Build a company brain](https://www.swfte.com/how-to-build-a-company-brain-for-ai): Extend this to many sources, identity-aware permissions and evidence on every answer.
- [Validate your AI](https://www.swfte.com/how-to-validate-your-ai): Turn the test set into regression gates and add human review.
- [Agentic RAG](https://www.swfte.com/agentic-rag): When one retrieval pass is not enough and the system needs to plan its searches.
- [Cost of RAG](https://www.swfte.com/cost-of-rag): Estimate embedding, storage and generation costs before you scale.
- [RAG vs MCP](https://www.swfte.com/compare/rag-vs-mcp): Decide between retrieving documents and letting an agent call tools.

## FAQ

### What is a RAG pipeline?

A RAG pipeline is the sequence that turns documents into answers: load and chunk the documents, embed the chunks, store them in an index, retrieve the best chunks for a question, and give those chunks to a language model to write a cited answer. This guide builds each stage in order.

### Can I run RAG locally without sending data to an API?

Yes. This guide runs the database, the embedding model and the chat model on your own machine. No document text leaves it. The trade-off is model size: a 3B local model writes weaker answers than a large hosted one, which is why retrieval quality matters more.

### Do I need a vector database?

Not necessarily. Postgres with pgvector stores vectors, text and permissions in one place and supports approximate indexes. A dedicated vector database makes sense when you need very large indexes or features Postgres lacks. Start with what your team already operates.

### What is hybrid search in RAG?

Hybrid search combines keyword matching with vector similarity and fuses the two rankings. Keywords catch exact terms and identifiers; vectors catch paraphrases. Reciprocal rank fusion, used here, adds 1 divided by (k plus rank) from each list, with k set to 60.

### How do I stop RAG leaking documents people should not see?

Store the access list on every chunk and filter inside the search query, before anything reaches the model. Filtering the answer afterwards is too late, because the model has already read the text. Add a test that must fail if the filter is removed.

### How do I know if my RAG system is any good?

Write real questions with the expected source file, then measure hit@3 and mean reciprocal rank for retrieval. Separately, read a sample of answers against their citations. Re-run both after every change to chunking, the embedding model or the prompt.

### When should I use fine-tuning instead of RAG?

Use RAG when the model needs facts from documents that change. Use fine-tuning when it needs a behaviour: a format, a tone or a narrow task. Many systems use both. Do not fine-tune to teach facts you will have to update next month.

## How Swfte can help

You can run every step above without Swfte. If you want retrieval inside a larger governed setup, these are the relevant parts.

- [Company brain](https://www.swfte.com/platform/company-brain): The design for permission-aware knowledge across many sources.
- [Cortex](https://www.swfte.com/products/cortex): A desktop app with knowledge bases and retrieval on the device, using local models through Ollama or LM Studio.
- [Swfte Connect](https://www.swfte.com/products/connect): An OpenAI-compatible gateway if you want to route the generation step to different models.

The company brain is a customer-hosted design. Parts are built, search over documents is in progress, and the details are on its page. Nothing in this guide depends on it.

## Sources

- [pgvector README (GitHub)](https://github.com/pgvector/pgvector): Docker image tag pgvector/pgvector:pg18-trixie, CREATE EXTENSION vector, vector columns, <=> cosine distance, HNSW index syntax with vector_cosine_ops, hnsw.iterative_scan values, filtering advice, hybrid search with RRF, version 0.8.7 and Postgres 13+
- [pgvector-python README (GitHub)](https://github.com/pgvector/pgvector-python): pip install pgvector, register_vector for psycopg 3, Vector type, example list including RAG with Ollama and hybrid search
- [pgvector-python RAG example (raw file)](https://raw.githubusercontent.com/pgvector/pgvector-python/master/examples/rag/example.py): ollama.embed with nomic-embed-text, search_document and search_query prefixes, vector(768), ollama.generate(...).response
- [PostgreSQL official Docker image](https://hub.docker.com/_/postgres): POSTGRES_PASSWORD, default port 5432, data volume path /var/lib/postgresql for PostgreSQL 18 and later
- [Python venv documentation](https://docs.python.org/3/library/venv.html): python -m venv and source <venv>/bin/activate
- [psycopg 3 installation](https://www.psycopg.org/psycopg3/docs/basic/install.html): pip install "psycopg[binary]", Python 3.10 to 3.15, PostgreSQL 10 to 18
- [Ollama quickstart](https://docs.ollama.com/quickstart): Download from ollama.com/download, local server at http://localhost:11434, ollama pull
- [Ollama embeddings](https://docs.ollama.com/capabilities/embeddings): /api/embed endpoint, list input for batches, L2-normalised vectors
- [nomic-embed-text on the Ollama library](https://ollama.com/library/nomic-embed-text): ollama pull nomic-embed-text, 274 MB, embedding only
- [llama3.2 on the Ollama library](https://ollama.com/library/llama3.2): Sizes 1b and 3b, latest 2.0 GB, 128K context
- [PostgreSQL text search functions](https://www.postgresql.org/docs/current/textsearch-controls.html): websearch_to_tsquery never raises syntax errors on user input, ts_rank_cd cover density ranking (PostgreSQL 18 docs)
- [PostgreSQL array operators](https://www.postgresql.org/docs/current/functions-array.html): The && overlap operator
- [PostgreSQL GIN index documentation](https://www.postgresql.org/docs/current/gin.html): GIN array_ops supports && and tsvector_ops supports @@
- [Sentence Transformers installation](https://sbert.net/docs/installation.html): pip install -U sentence-transformers, recommended Python 3.10+ and PyTorch 2.2+
- [Sentence Transformers CrossEncoder usage](https://sbert.net/docs/cross_encoder/usage/usage.html): CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2") and model.rank(query, passages) returning corpus_id and score
- [cross-encoder/ms-marco-MiniLM-L6-v2 model card](https://huggingface.co/cross-encoder/ms-marco-MiniLM-L6-v2): 22.7M parameters, Apache 2.0 licence

Last verified against these sources on 2026-10-06.
