Large language models know a lot, but they don’t know your stuff: your company wiki, your product manuals or the 200 PDFs in your research folder. Retrieval-augmented generation (RAG) fixes that by fetching the relevant pieces of your documents and handing them to the model along with the question.
In this tutorial we’ll build a tiny but complete RAG pipeline in Python.
How RAG works in 4 steps
- Chunk — split your documents into small passages (a few hundred words each)
- Embed — turn every chunk into a vector that captures its meaning
- Retrieve — when a question comes in, embed it too and find the most similar chunks
- Generate — send the question plus those chunks to an LLM and ask it to answer using only that context
The magic is in step 3: you’re searching by meaning, not keywords. A question about “time off” will find a paragraph about “annual leave policy” even though no words match.
Setting up
Create a project folder and install a couple of packages:
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install sentence-transformers numpy anthropic
We’ll use sentence-transformers to create embeddings locally (free, no API key) and an LLM API for the final answer.
Step 1: Load and chunk documents
from pathlib import Path
def load_chunks(folder: str, size: int = 120) -> list[str]:
chunks = []
for file in Path(folder).glob("*.txt"):
words = file.read_text(encoding="utf-8").split()
for i in range(0, len(words), size):
chunks.append(" ".join(words[i:i + size]))
return chunks
chunks = load_chunks("docs")
print(f"Loaded {len(chunks)} chunks")
Overlapping chunks (e.g. 120 words with a 20-word overlap) often improve results, because an important sentence won’t get cut in half at a boundary.
Step 2: Embed every chunk
from sentence_transformers import SentenceTransformer
import numpy as np
model = SentenceTransformer("all-MiniLM-L6-v2")
vectors = model.encode(chunks, normalize_embeddings=True)
Each chunk is now a 384-number vector. Because we normalised them, we can measure similarity with a simple dot product.
Step 3: Retrieve the best matches
def search(question: str, k: int = 4) -> list[str]:
q = model.encode([question], normalize_embeddings=True)[0]
scores = vectors @ q
best = np.argsort(scores)[::-1][:k]
return [chunks[i] for i in best]
That’s a working semantic search engine in five lines. For thousands of documents you’d swap the NumPy array for a vector database such as pgvector, Qdrant or Chroma — but the idea is identical.
Step 4: Generate the answer
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY from env
def ask(question: str) -> str:
context = "\n\n---\n\n".join(search(question))
prompt = (
"Answer the question using only the context below. "
"If the answer isn't there, say you don't know.\n\n"
f"Context:\n{context}\n\nQuestion: {question}"
)
reply = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=600,
messages=[{"role": "user", "content": prompt}],
)
return reply.content[0].text
print(ask("How many days of annual leave do new employees get?"))
The instruction “If the answer isn’t there, say you don’t know” is important — it dramatically reduces made-up answers.
Making it production-ready
Our toy version works, but real apps usually add:
- Better chunking — split on headings and paragraphs instead of fixed word counts
- Metadata — store the source file and page so you can show citations
- Hybrid search — combine semantic search with classic keyword search (BM25) for names, codes and numbers
- Re-ranking — fetch 20 candidates, then use a smaller model to pick the best 4
- Evaluation — keep a list of real questions with known answers and test every change against it
Never put confidential documents into a third-party API without checking your organisation’s data policy. For sensitive data, look at self-hosted open-source models.
Wrapping up
RAG is the most practical way to make AI useful with your own information. You don’t need to train a model — just organise your documents, embed them and retrieve wisely. Start with the 40-line version above, point it at a folder of notes and you’ll be surprised how capable it is.
Built this over the weekend for our company handbook. Works surprisingly well with just 40 lines!