This project implements a Retrieval-Augmented Generation (RAG) pipeline to help users manage and retrieve information about 2025–2026 concert tours. It was built as part of the Provectus Internship 2025 selection task.
Build a system that:
- Ingests and summarizes concert-related documents
- Answers user questions grounded only in those documents
- Falls back to a live web search if the answer is not available locally
-
Summarization: Used the
facebook/bart-large-cnnmodel via Hugging Face. Since BART sometimes omits key names, a named-entity patching step was added to ensure artists like "Lady Gaga" are retained in summaries. -
Relevance Filtering: Lightweight keyword-matching function checks if a document is concert-related before ingestion.
-
Storage & Retrieval: All summaries are embedded using
sentence-transformers/bert-base-nli-mean-tokensand stored in a FAISS vector index. -
Question Answering: Used a pre-trained
bert-large-uncased-whole-word-masking-finetuned-squadmodel to extract answers from the retrieved summaries. spaCy is used to extract artist names from the user's question and filter results accordingly. -
Web Search Fallback: Integrated SerpAPI for real-time Google search results when no relevant documents are found.
-
Interface: Simple command-line interface (CLI) with 4 options:
1. Ingest a document (paste text)
2. Ask a question from database
3. Web search (online)
4. Exit
ProvectusInternship_ValentinaMkrtumyan/
├── main.py # CLI entry point
├── requirements.txt # All required dependencies
├── .env # Stores SERPAPI key (excluded from Git)
├── ingestion/
│ ├── summarizer.py # Summarizes concert documents
│ ├── relevance_checker.py # Checks if doc is concert-related
│ └── document_processor.py # Handles ingestion logic
├── rag_engine/
│ ├── vector_store.py # Embedding + FAISS index
│ └── query_answering.py # Answers questions using retrieved docs
├── websearch/
│ └── search_fallback.py # Web fallback using SerpAPI
├── data/
│ ├── documents.json # Saved document summaries
│ └── faiss_index/
| │ ├── index.faiss # FAISS vector index
| │ └── docs.json # Stored vector-related summaries
- Clone the repo
git clone https://github.com/valentina_mkrtu/ProvectusInternship_ValentinaMkrtumyan.git
cd ProvectusInternship_ValentinaMkrtumyan- Create and activate a virtual environment
python -m venv venv
venv\Scripts\activate # Windows- Install dependencies
pip install -r requirements.txt
python -m spacy download en_core_web_sm- Set your SerpAPI key in .env
# .env
SERPAPI_API_KEY=your_serpapi_key_herepython main.py🎤 Concert Tour Assistant - CLI
- Ingest a document (paste text)
- Ask a question from database
- Web search (online)
- Exit
Example workflow:
Paste a concert tour announcement → it gets summarized + indexed
Ask a question → answered from indexed documents
Ask about an unknown artist → fallback to Google (via SerpAPI)
Example Usage:
Option 1: Paste a concert tour document about an artist (e.g., Lady Gaga) → System checks relevance → summarizes → stores it
Option 2: Ask "When is Lady Gaga performing?" → Answer retrieved from summarized & stored data
Option 3: Ask about an artist not in the documents (e.g., "Where is Chappell Roan performing?") → System searches the web using SerpAPI and returns top results
Transformers (BART, BERT from HuggingFace)
FAISS (vector similarity search)
Sentence-BERT (bert-base-nli-mean-tokens)
spaCy (NER and question analysis)
SerpAPI (Google fallback for artist info)
Python CLI (no frontend UI for this version)
The .env file with API keys is ignored via .gitignore
No API keys or sensitive files are tracked
Valentina Mkrtumyan Data Science & AI Enthusiast 📍 Based in Armenia