Skip to content
Nexoora Club
AI & Machine Learning

Build Your First RAG App in Python: Chat With Your Own Documents

Retrieval-augmented generation lets an AI answer questions using your PDFs, notes and docs. Here's how it works and a minimal app you can build in an afternoon.

nexooraclub3 min read1

Large language models know a lot, but they don’t know your stuff: your company wiki, your product manuals or the 200 PDFs in your research folder. Retrieval-augmented generation (RAG) fixes that by fetching the relevant pieces of your documents and handing them to the model along with the question.

In this tutorial we’ll build a tiny but complete RAG pipeline in Python.

How RAG works in 4 steps

  1. Chunk — split your documents into small passages (a few hundred words each)
  2. Embed — turn every chunk into a vector that captures its meaning
  3. Retrieve — when a question comes in, embed it too and find the most similar chunks
  4. Generate — send the question plus those chunks to an LLM and ask it to answer using only that context

The magic is in step 3: you’re searching by meaning, not keywords. A question about “time off” will find a paragraph about “annual leave policy” even though no words match.

Setting up

Create a project folder and install a couple of packages:

python -m venv .venv
source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install sentence-transformers numpy anthropic

We’ll use sentence-transformers to create embeddings locally (free, no API key) and an LLM API for the final answer.

Step 1: Load and chunk documents

from pathlib import Path

def load_chunks(folder: str, size: int = 120) -> list[str]:
    chunks = []
    for file in Path(folder).glob("*.txt"):
        words = file.read_text(encoding="utf-8").split()
        for i in range(0, len(words), size):
            chunks.append(" ".join(words[i:i + size]))
    return chunks

chunks = load_chunks("docs")
print(f"Loaded {len(chunks)} chunks")

Overlapping chunks (e.g. 120 words with a 20-word overlap) often improve results, because an important sentence won’t get cut in half at a boundary.

Step 2: Embed every chunk

from sentence_transformers import SentenceTransformer
import numpy as np

model = SentenceTransformer("all-MiniLM-L6-v2")
vectors = model.encode(chunks, normalize_embeddings=True)

Each chunk is now a 384-number vector. Because we normalised them, we can measure similarity with a simple dot product.

Step 3: Retrieve the best matches

def search(question: str, k: int = 4) -> list[str]:
    q = model.encode([question], normalize_embeddings=True)[0]
    scores = vectors @ q
    best = np.argsort(scores)[::-1][:k]
    return [chunks[i] for i in best]

That’s a working semantic search engine in five lines. For thousands of documents you’d swap the NumPy array for a vector database such as pgvector, Qdrant or Chroma — but the idea is identical.

Step 4: Generate the answer

import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY from env

def ask(question: str) -> str:
    context = "\n\n---\n\n".join(search(question))
    prompt = (
        "Answer the question using only the context below. "
        "If the answer isn't there, say you don't know.\n\n"
        f"Context:\n{context}\n\nQuestion: {question}"
    )
    reply = client.messages.create(
        model="claude-sonnet-5-5",
        max_tokens=600,
        messages=[{"role": "user", "content": prompt}],
    )
    return reply.content[0].text

print(ask("How many days of annual leave do new employees get?"))

The instruction “If the answer isn’t there, say you don’t know” is important — it dramatically reduces made-up answers.

Making it production-ready

Our toy version works, but real apps usually add:

  • Better chunking — split on headings and paragraphs instead of fixed word counts
  • Metadata — store the source file and page so you can show citations
  • Hybrid search — combine semantic search with classic keyword search (BM25) for names, codes and numbers
  • Re-ranking — fetch 20 candidates, then use a smaller model to pick the best 4
  • Evaluation — keep a list of real questions with known answers and test every change against it

Never put confidential documents into a third-party API without checking your organisation’s data policy. For sensitive data, look at self-hosted open-source models.

Wrapping up

RAG is the most practical way to make AI useful with your own information. You don’t need to train a model — just organise your documents, embed them and retrieve wisely. Start with the 40-line version above, point it at a folder of notes and you’ll be surprised how capable it is.

Written by

nexooraclub

Writer and builder covering software, AI and the tools that make developers faster. Add your bio under Users → Profile.

1 comment

Join the discussion

Your email address will not be published. Required fields are marked *