Skip to content

Latest commit

 

History

17 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🎤 Concert Tour Assistant — RAG System

This project implements a Retrieval-Augmented Generation (RAG) pipeline to help users manage and retrieve information about 2025–2026 concert tours. It was built as part of the Provectus Internship 2025 selection task.


Approach & Design Choices

Objective:

Build a system that:

  • Ingests and summarizes concert-related documents
  • Answers user questions grounded only in those documents
  • Falls back to a live web search if the answer is not available locally

Key Design Decisions:

  • Summarization: Used the facebook/bart-large-cnn model via Hugging Face. Since BART sometimes omits key names, a named-entity patching step was added to ensure artists like "Lady Gaga" are retained in summaries.

  • Relevance Filtering: Lightweight keyword-matching function checks if a document is concert-related before ingestion.

  • Storage & Retrieval: All summaries are embedded using sentence-transformers/bert-base-nli-mean-tokens and stored in a FAISS vector index.

  • Question Answering: Used a pre-trained bert-large-uncased-whole-word-masking-finetuned-squad model to extract answers from the retrieved summaries. spaCy is used to extract artist names from the user's question and filter results accordingly.

  • Web Search Fallback: Integrated SerpAPI for real-time Google search results when no relevant documents are found.

  • Interface: Simple command-line interface (CLI) with 4 options:

  1. Ingest a document (paste text)
  2. Ask a question from database
  3. Web search (online)
  4. Exit

Project Structure

ProvectusInternship_ValentinaMkrtumyan/
├── main.py # CLI entry point
├── requirements.txt # All required dependencies
├── .env # Stores SERPAPI key (excluded from Git) 
├── ingestion/ 
│ ├── summarizer.py # Summarizes concert documents 
│ ├── relevance_checker.py # Checks if doc is concert-related 
│ └── document_processor.py # Handles ingestion logic
├── rag_engine/ 
│ ├── vector_store.py # Embedding + FAISS index 
│ └── query_answering.py # Answers questions using retrieved docs
├── websearch/ 
│ └── search_fallback.py # Web fallback using SerpAPI
├── data/ 
│ ├── documents.json # Saved document summaries 
│ └── faiss_index/ 
| │ ├── index.faiss # FAISS vector index 
| │ └── docs.json # Stored vector-related summaries

⚙️ Setup Instructions

  1. Clone the repo
git clone https://github.com/valentina_mkrtu/ProvectusInternship_ValentinaMkrtumyan.git
cd ProvectusInternship_ValentinaMkrtumyan
  1. Create and activate a virtual environment
python -m venv venv
venv\Scripts\activate   # Windows
  1. Install dependencies
pip install -r requirements.txt
python -m spacy download en_core_web_sm
  1. Set your SerpAPI key in .env
# .env
SERPAPI_API_KEY=your_serpapi_key_here

🖥️ How to Use (CLI)

python main.py

You’ll see the menu:

🎤 Concert Tour Assistant - CLI

  1. Ingest a document (paste text)
  2. Ask a question from database
  3. Web search (online)
  4. Exit

Example workflow:

Paste a concert tour announcement → it gets summarized + indexed

Ask a question → answered from indexed documents

Ask about an unknown artist → fallback to Google (via SerpAPI)

Example Usage:

Option 1: Paste a concert tour document about an artist (e.g., Lady Gaga) → System checks relevance → summarizes → stores it

Option 2: Ask "When is Lady Gaga performing?" → Answer retrieved from summarized & stored data

Option 3: Ask about an artist not in the documents (e.g., "Where is Chappell Roan performing?") → System searches the web using SerpAPI and returns top results

🧠 Technologies Used

Transformers (BART, BERT from HuggingFace)

FAISS (vector similarity search)

Sentence-BERT (bert-base-nli-mean-tokens)

spaCy (NER and question analysis)

SerpAPI (Google fallback for artist info)

Python CLI (no frontend UI for this version)

🔒 Security

The .env file with API keys is ignored via .gitignore

No API keys or sensitive files are tracked

🧑‍💻 Author

Valentina Mkrtumyan Data Science & AI Enthusiast 📍 Based in Armenia

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages