Skip to main content
AI / Machine Learning

RAG System Implementation

10 Modules
Chapter 8: Resources: RAG System Implementation80%

Resources: RAG System Implementation

Introduction

This resource collection supports your ongoing RAG development journey. It covers official documentation, community tools, production deployment guides, and further reading -- organised by topic and skill level.


Core Libraries and Documentation

Vector Databases

ChromaDB (Local-First)

Qdrant (Production-Grade)

Pinecone (Managed Cloud)

Weaviate (Hybrid + GraphQL)

pgvector (PostgreSQL Extension)

RAG Frameworks

LangChain

LlamaIndex

Haystack (by deepset)

Embedding Models

OpenAI Embeddings

Sentence Transformers (Local, Free)

Ollama Embeddings (Local, GPU-Accelerated)

Cohere Embed

Voyage AI

  • Documentation: docs.voyageai.com
  • Models: voyage-3, voyage-code-3
  • Strong performance on retrieval benchmarks

Evaluation and Quality

RAGAS (RAG Assessment)

The standard framework for evaluating RAG pipeline quality.

# Quick RAGAS evaluation example
from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy, context_precision

result = evaluate(
    dataset=your_test_dataset,
    metrics=[faithfulness, answer_relevancy, context_precision],
)
print(result)

DeepEval

TruLens


Re-ranking Models

Cross-Encoders

  • cross-encoder/ms-marco-MiniLM-L-6-v2 -- Fast, good quality
  • cross-encoder/ms-marco-MiniLM-L-12-v2 -- Better quality, slightly slower
  • BAAI/bge-reranker-base -- Strong retrieval-focused reranker
  • BAAI/bge-reranker-large -- Best quality reranker
# Install
pip install sentence-transformers

Cohere Rerank


Document Processing

PDF Extraction

OCR (Scanned Documents)

Structured Data

  • LlamaIndex supports CSV, JSON, SQL, and API data sources
  • Pandas + LangChain for tabular data integration
  • SQLAlchemy for database-backed RAG systems

Local AI Infrastructure

Ollama

  • Website: ollama.com
  • Model Library: ollama.com/library
  • Recommended models for RAG:
    • Generation: llama3.3:8b, mistral, qwen2.5:7b
    • Embeddings: nomic-embed-text, mxbai-embed-large
    • Code understanding: qwen2.5-coder:7b

LM Studio

  • Website: lmstudio.ai
  • GUI-based model management and inference
  • OpenAI-compatible local server
  • Good for experimentation and non-technical users

vLLM (Production Serving)


Advanced RAG Patterns

Agentic RAG

Combine RAG with tool-using agents for multi-step reasoning:

Multi-Modal RAG

Retrieve and reason over images, tables, and text together:

Graph RAG

Combine knowledge graphs with vector retrieval:

Corrective RAG (CRAG)

Self-correcting retrieval that validates and re-fetches when results are poor:

  • Paper: "Corrective Retrieval Augmented Generation" (Yan et al., 2024)
  • Combines retrieval evaluation with web search fallback

Deployment and Production

Web Interfaces

Streamlit (Quick prototyping)

Gradio (ML-focused UI)

  • Documentation: gradio.app/docs
  • Installation: pip install gradio
  • Chatbot component for conversational RAG

Chainlit (Chat-first UI for LLM apps)

  • Documentation: docs.chainlit.io
  • Built specifically for conversational AI applications

API Deployment

FastAPI

  • Documentation: fastapi.tiangolo.com
  • Ideal for building RAG-as-a-service APIs
  • Async support for high concurrency
# Minimal FastAPI RAG endpoint
from fastapi import FastAPI
from rag_pipeline import ask

app = FastAPI()

@app.post("/ask")
async def ask_question(query: str):
    result = ask(query)
    return result

Docker Deployment

# Dockerfile for RAG service
FROM python:3.11-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install -r requirements.txt
COPY scripts/ scripts/
COPY data/ data/
CMD ["uvicorn", "scripts.api:app", "--host", "0.0.0.0", "--port", "8000"]

Recommended Reading

Foundational Papers

  1. "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks" (Lewis et al., 2020) The original RAG paper from Meta AI. Essential reading for understanding the paradigm.

  2. "Dense Passage Retrieval for Open-Domain Question Answering" (Karpukhin et al., 2020) Foundational work on using dense embeddings for document retrieval.

  3. "Attention Is All You Need" (Vaswani et al., 2017) The Transformer paper -- understanding this helps you grasp how embeddings and LLMs work.

  4. "REALM: Retrieval-Augmented Language Model Pre-Training" (Guu et al., 2020) Integrating retrieval directly into the pre-training process.

Practical Guides

  1. "Building RAG-based LLM Applications for Production" -- Anyscale Blog Comprehensive guide to moving from prototype to production.

  2. "A Survey on Retrieval-Augmented Generation" (Gao et al., 2024) Thorough survey of RAG techniques, evaluation, and future directions.

  3. "Chunking Strategies for LLM Applications" -- Pinecone Learning Centre Detailed comparison of chunking approaches with benchmarks.

Books

  1. "Building LLM Apps" by Valentina Alto (O'Reilly, 2024) Practical guide covering RAG, agents, and deployment.

  2. "Designing Machine Learning Systems" by Chip Huyen (O'Reilly, 2022) Production ML systems design -- relevant to RAG infrastructure.


Community and Support

Forums and Communities

  • LangChain Discord: Active community for LangChain + RAG questions
  • LlamaIndex Discord: Data framework discussions and support
  • Chroma Discord: ChromaDB-specific help and feature discussions
  • r/LocalLLaMA (Reddit): Local model running and RAG discussions
  • Hugging Face Forums: Model selection and embedding questions

Newsletters and Blogs

  • The Batch (deeplearning.ai): Weekly AI news including RAG developments
  • LangChain Blog: blog.langchain.dev -- Tutorials and patterns
  • Pinecone Learning Centre: pinecone.io/learn -- Vector search education
  • Weaviate Blog: weaviate.io/blog -- Hybrid search and RAG patterns

Video Courses

  • "LangChain for LLM Application Development" -- DeepLearning.AI (short course, free)
  • "Building and Evaluating Advanced RAG Applications" -- DeepLearning.AI (short course, free)
  • "Vector Databases: from Embeddings to Applications" -- DeepLearning.AI (short course, free)

Quick Reference Card

Essential Commands

# Environment setup
python -m venv venv && source venv/bin/activate
pip install chromadb langchain sentence-transformers ollama

# Ollama model management
ollama pull llama3.3:8b          # Download generation model
ollama pull nomic-embed-text     # Download embedding model
ollama list                      # List installed models
ollama serve                     # Start server

# Project scripts (from ~/rag-workshop)
python scripts/verify_setup.py         # Check installation
python scripts/build_index.py          # Build vector store
python scripts/rag_pipeline.py --chat  # Interactive Q&A
python scripts/add_document.py <file>  # Add new document
python scripts/evaluate_rag.py         # Run quality tests

Key Python Imports

# Vector store
import chromadb
from chromadb.utils import embedding_functions

# Document processing
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain.docstore.document import Document

# Local embeddings
from sentence_transformers import SentenceTransformer

# Re-ranking
from sentence_transformers import CrossEncoder

# Local LLM
import ollama

# Cloud embeddings (optional)
from openai import OpenAI

Embedding Model Quick Selection

Use CaseModelCommand
Learning / Prototypingall-MiniLM-L6-v2pip install sentence-transformers
Local Productionnomic-embed-textollama pull nomic-embed-text
Cloud (Best Quality)text-embedding-3-smallOpenAI API key required
MultilingualBAAI/bge-m3pip install sentence-transformers

Navigation