An end-to-end Retrieval-Augmented Generation (RAG) knowledge base QA system built with LangChain, BGE-M3, FAISS, BM25, RRF, BGE Reranker, and Ollama.
The project implements and compares multiple retrieval strategies to study how different retrieval components affect the final RAG answer quality.
Documents
│
▼
Document Loading
│
▼
Chunking
│
┌──────────┴──────────┐
▼ ▼
BGE-M3 Embedding BM25
│ │
▼ ▼
FAISS Lexical Retrieval
│ │
└──────────┬──────────┘
▼
Hybrid Retrieval
│
RRF
│
▼
Candidate Chunks
│
▼
BGE Reranker (Optional)
│
▼
Top-K Chunks
│
▼
Local LLM
Ollama
│
▼
Answer
- PDF / TXT / Markdown document loading
- Recursive text chunking
- Dense retrieval with BGE-M3
- Vector search with FAISS
- Lexical retrieval with BM25
- Hybrid retrieval using Reciprocal Rank Fusion (RRF)
- Second-stage reranking with BGE-Reranker-v2-M3
- Local LLM generation with Ollama
- Retrieval method comparison
- Retrieval evaluation with Recall@K and MRR@K
- Generation evaluation for answer correctness and faithfulness
The project supports four retrieval configurations:
| Method | Pipeline |
|---|---|
| FAISS | FAISS → Top-K → LLM |
| BM25 | BM25 → Top-K → LLM |
| Hybrid | FAISS + BM25 → RRF → Top-K → LLM |
| Hybrid + Reranker | FAISS + BM25 → RRF → Reranker → Top-K → LLM |
The retrieval method can be selected in run_rag.py:
# 0 - FAISS
# 1 - BM25
# 2 - Hybrid
# 3 - Hybrid + Reranker
RETRIEVAL_METHOD = 0Dense retrieval and lexical retrieval have different strengths.
Dense retrieval
BGE-M3 converts documents and queries into embeddings and retrieves semantically similar chunks.
It is useful when the query and document use different wording but have similar meanings.
Query:
"What is the purpose of vector embeddings?"
Document:
"Embeddings represent text as dense numerical vectors..."
BM25
BM25 focuses on lexical matching and term importance.
It can perform well when the query contains important technical terms, names, or exact keywords.
Hybrid Retrieval
The project combines both approaches:
FAISS
│
├── semantic relevance
│
▼
Top-K documents
\
\
→ RRF → Final ranking
/
/
BM25
│
└── lexical relevance
FAISS and BM25 produce scores with different scales, so their raw scores are not directly comparable.
Instead, this project uses Reciprocal Rank Fusion (RRF).
For a document with rank r:
RRF(d) = 1 / (k + r)
where k is typically set to 60.
If a document appears in both retrieval results, its scores are accumulated:
RRF(d) =
1 / (60 + rank_FAISS)
+ 1 / (60 + rank_BM25)
This allows the system to combine multiple ranked lists without requiring their original scores to be on the same scale.
The hybrid retriever first retrieves a larger candidate set:
FAISS Top-8
+
BM25 Top-8
↓
RRF
↓
Hybrid candidates
↓
BGE-Reranker-v2-M3
↓
Top-4
↓
LLM
The reranker uses a cross-encoder to directly evaluate the relevance between:
(query, document)
This is different from the embedding-based retrieval stage.
The general design is:
First-stage retrieval
↓
High recall
↓
Candidate documents
↓
Cross-encoder reranking
↓
High precision
↓
LLM
langchain-rag-knowledge-base/
│
├── data/
│ └── documents/
│ └── *.pdf / *.txt / *.md
│
├── src/
│ ├── __init__.py
│ ├── ingestion.py
│ ├── embeddings.py
│ ├── vector_store.py
│ ├── bm25.py
│ ├── hybrid.py
│ ├── reranker.py
│ └── generation.py
│
├── evaluation/
│ └── questions.json
│
├── tests/
│
├── run_rag.py
├── requirements.txt
├── .env.example
├── .gitignore
└── README.md
python -m venv .venvActivate it on Windows:
.venv\Scripts\activatepip install -r requirements.txtInstall Ollama and pull a local LLM:
ollama pull qwen2.5:1.5bStart Ollama if it is not already running:
ollama serveThe default Ollama endpoint is:
http://localhost:11434
The project uses:
BAAI/bge-m3
BAAI/bge-reranker-v2-m3
Qwen 2.5 1.5B
through Ollama.
If Hugging Face access is unavailable, a mirror can be configured:
export HF_ENDPOINT=https://hf-mirror.comOn Windows PowerShell:
$env:HF_ENDPOINT="https://hf-mirror.com"The Python process will then use the configured Hugging Face endpoint when downloading models.
Place knowledge-base documents inside:
data/documents/
For example:
data/documents/
├── rag_basics.md
├── vector_database.md
├── bm25.md
├── embedding.md
├── reranking.md
└── transformer.md
The current ingestion pipeline supports:
.pdf
.txt
.md
Run:
python run_rag.pyThe program will:
1. Load documents
2. Split documents into chunks
3. Load/build the FAISS index
4. Initialize BM25
5. Initialize the hybrid retriever
6. Initialize the reranker
7. Retrieve relevant chunks
8. Generate an answer with the local LLM
Then enter a question:
Please input your query:
What is BM25?
Type:
q
to exit.
The main retrieval configuration is controlled by:
RETRIEVAL_METHOD = 0
RETRIEVAL_K = 8
RERANK_TOP_K = 4FAISS → Top-4 → LLM
BM25 → Top-4 → LLM
FAISS + BM25
↓
RRF
↓
Top-4
↓
LLM
FAISS + BM25
↓
RRF
↓
Top-8
↓
BGE Reranker
↓
Top-4
↓
LLM
Using the same number of final chunks for the four methods makes the comparison of final LLM answers more controlled.
The project evaluates the retrieval pipeline independently from the generation pipeline.
Two primary metrics are used:
Measures whether at least one relevant document/chunk appears in the top K results.
Recall@K =
# queries with relevant result in Top-K
---------------------------------------
# total queries
Measures how highly the first relevant result is ranked.
MRR = 1 / rank_of_first_relevant_document
The retrieval experiment can compare:
| Method | Recall@5 | MRR@5 |
|---|---|---|
| FAISS | - | - |
| BM25 | - | - |
| Hybrid | - | - |
| Hybrid + Reranker | - | - |
The goal is to determine whether each retrieval component improves retrieval quality.
Generation quality can be evaluated using:
- Answer correctness
- Faithfulness
- Context relevance
- LLM-as-a-Judge
A useful experiment is:
FAISS
vs
BM25
vs
Hybrid
vs
Hybrid + Reranker
while keeping the same:
- Questions
- Knowledge base
- LLM
- Prompt
- Number of final context chunks
This makes the comparison more meaningful.
The evaluation questions should be constructed from the project's own knowledge base.
Example:
[
{
"id": 1,
"question": "What is BM25?",
"relevant_chunks": [37],
"reference_answer": "BM25 is a classical lexical information retrieval algorithm."
}
]Questions should include:
- Exact keyword queries
- Semantic/paraphrased queries
- Queries requiring hybrid retrieval
- Queries where multiple chunks are similar
- Unanswerable questions
This helps evaluate the strengths and weaknesses of different retrieval strategies.
A typical experiment can use:
6 documents
↓
86 chunks
↓
┌──────────────────────────────┐
│ │
FAISS BM25
│ │
└──────────────┬───────────────┘
│
RRF
│
Hybrid Retrieval
│
BGE Reranker
│
Top-4
│
Qwen 2.5
│
Answer
The experiment can then measure how retrieval quality changes when adding:
FAISS
↓
BM25
↓
Hybrid / RRF
↓
Reranker
Python
LangChain
BGE-M3
BGE-Reranker-v2-M3
FAISS
BM25
RRF
Ollama
Qwen 2.5
This project demonstrates the following RAG concepts:
Document Processing
↓
Chunking
↓
Embedding
↓
Dense Retrieval
↓
Sparse Retrieval
↓
Hybrid Retrieval
↓
Rank Fusion
↓
Reranking
↓
Context Construction
↓
LLM Generation
↓
Evaluation
The project therefore covers the major components of a modern retrieval-augmented generation system from document ingestion to final answer evaluation.