{"id":108,"date":"2026-09-29T11:00:00","date_gmt":"2026-09-29T06:00:00","guid":{"rendered":"https:\/\/example.com\/?p=108"},"modified":"2026-09-29T11:00:00","modified_gmt":"2026-09-29T06:00:00","slug":"build-your-first-rag-app-python","status":"publish","type":"post","link":"https:\/\/www.nexooraclub.com\/?p=108","title":{"rendered":"Build Your First RAG App in Python: Chat With Your Own Documents"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Large language models know a lot, but they don&#8217;t know <em>your<\/em> stuff: your company wiki, your product manuals or the 200 PDFs in your research folder. <strong>Retrieval-augmented generation (RAG)<\/strong> fixes that by fetching the relevant pieces of your documents and handing them to the model along with the question.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In this tutorial we&#8217;ll build a tiny but complete RAG pipeline in Python.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">How RAG works in 4 steps<\/h2>\n\n\n\n<ol class=\"wp-block-list\">\n<li><strong>Chunk<\/strong> \u2014 split your documents into small passages (a few hundred words each)<\/li>\n\n\n<li><strong>Embed<\/strong> \u2014 turn every chunk into a vector that captures its meaning<\/li>\n\n\n<li><strong>Retrieve<\/strong> \u2014 when a question comes in, embed it too and find the most similar chunks<\/li>\n\n\n<li><strong>Generate<\/strong> \u2014 send the question plus those chunks to an LLM and ask it to answer <em>using only that context<\/em><\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">The magic is in step 3: you&#8217;re searching by <strong>meaning<\/strong>, not keywords. A question about &#8220;time off&#8221; will find a paragraph about &#8220;annual leave policy&#8221; even though no words match.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Setting up<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Create a project folder and install a couple of packages:<\/p>\n\n\n\n<pre class=\"wp-block-code language-bash\"><code>python -m venv .venv\nsource .venv\/bin\/activate   # Windows: .venv\\Scripts\\activate\npip install sentence-transformers numpy anthropic<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">We&#8217;ll use <code>sentence-transformers<\/code> to create embeddings locally (free, no API key) and an LLM API for the final answer.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Step 1: Load and chunk documents<\/h2>\n\n\n\n<pre class=\"wp-block-code language-python\"><code>from pathlib import Path\n\ndef load_chunks(folder: str, size: int = 120) -&gt; list[str]:\n    chunks = []\n    for file in Path(folder).glob(\"*.txt\"):\n        words = file.read_text(encoding=\"utf-8\").split()\n        for i in range(0, len(words), size):\n            chunks.append(\" \".join(words[i:i + size]))\n    return chunks\n\nchunks = load_chunks(\"docs\")\nprint(f\"Loaded {len(chunks)} chunks\")<\/code><\/pre>\n\n\n\n<p class=\"tp-callout wp-block-paragraph\">Overlapping chunks (e.g. 120 words with a 20-word overlap) often improve results, because an important sentence won&#8217;t get cut in half at a boundary.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Step 2: Embed every chunk<\/h2>\n\n\n\n<pre class=\"wp-block-code language-python\"><code>from sentence_transformers import SentenceTransformer\nimport numpy as np\n\nmodel = SentenceTransformer(\"all-MiniLM-L6-v2\")\nvectors = model.encode(chunks, normalize_embeddings=True)<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Each chunk is now a 384-number vector. Because we normalised them, we can measure similarity with a simple dot product.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Step 3: Retrieve the best matches<\/h2>\n\n\n\n<pre class=\"wp-block-code language-python\"><code>def search(question: str, k: int = 4) -&gt; list[str]:\n    q = model.encode([question], normalize_embeddings=True)[0]\n    scores = vectors @ q\n    best = np.argsort(scores)[::-1][:k]\n    return [chunks[i] for i in best]<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">That&#8217;s a working semantic search engine in five lines. For thousands of documents you&#8217;d swap the NumPy array for a vector database such as pgvector, Qdrant or Chroma \u2014 but the idea is identical.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Step 4: Generate the answer<\/h2>\n\n\n\n<pre class=\"wp-block-code language-python\"><code>import anthropic\n\nclient = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY from env\n\ndef ask(question: str) -&gt; str:\n    context = \"\\n\\n---\\n\\n\".join(search(question))\n    prompt = (\n        \"Answer the question using only the context below. \"\n        \"If the answer isn't there, say you don't know.\\n\\n\"\n        f\"Context:\\n{context}\\n\\nQuestion: {question}\"\n    )\n    reply = client.messages.create(\n        model=\"claude-sonnet-5-5\",\n        max_tokens=600,\n        messages=[{\"role\": \"user\", \"content\": prompt}],\n    )\n    return reply.content[0].text\n\nprint(ask(\"How many days of annual leave do new employees get?\"))<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">The instruction <em>&#8220;If the answer isn&#8217;t there, say you don&#8217;t know&#8221;<\/em> is important \u2014 it dramatically reduces made-up answers.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Making it production-ready<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Our toy version works, but real apps usually add:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Better chunking<\/strong> \u2014 split on headings and paragraphs instead of fixed word counts<\/li>\n\n\n<li><strong>Metadata<\/strong> \u2014 store the source file and page so you can show citations<\/li>\n\n\n<li><strong>Hybrid search<\/strong> \u2014 combine semantic search with classic keyword search (BM25) for names, codes and numbers<\/li>\n\n\n<li><strong>Re-ranking<\/strong> \u2014 fetch 20 candidates, then use a smaller model to pick the best 4<\/li>\n\n\n<li><strong>Evaluation<\/strong> \u2014 keep a list of real questions with known answers and test every change against it<\/li>\n<\/ul>\n\n\n\n<p class=\"tp-callout tp-callout--warn wp-block-paragraph\">Never put confidential documents into a third-party API without checking your organisation&#8217;s data policy. For sensitive data, look at self-hosted open-source models.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Wrapping up<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">RAG is the most practical way to make AI useful with your own information. You don&#8217;t need to train a model \u2014 just organise your documents, embed them and retrieve wisely. Start with the 40-line version above, point it at a folder of notes and you&#8217;ll be surprised how capable it is.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Retrieval-augmented generation lets an AI answer questions using your PDFs, notes and docs. Here&#8217;s how it works and a minimal app you can build in an afternoon.<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[9,30,31,34],"class_list":["post-108","post","type-post","status-publish","format-standard","hentry","category-ai-machine-learning","tag-ai","tag-python","tag-rag","tag-tutorial"],"_links":{"self":[{"href":"https:\/\/www.nexooraclub.com\/index.php?rest_route=\/wp\/v2\/posts\/108","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.nexooraclub.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.nexooraclub.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.nexooraclub.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.nexooraclub.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=108"}],"version-history":[{"count":0,"href":"https:\/\/www.nexooraclub.com\/index.php?rest_route=\/wp\/v2\/posts\/108\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.nexooraclub.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=108"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.nexooraclub.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=108"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.nexooraclub.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=108"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}