- status: active
- type: explanation
- description: Overview, setup guide, and technical architecture reference for the MCMP Chatbot application.
A structured-data chatbot for the Munich Center for Mathematical Philosophy (MCMP). This application scrapes the MCMP website for the latest events, people, and research, and uses an LLM (Google Gemini) with structured MCP tools to answer user queries about the center's activities.
The application is built with Streamlit for the frontend, uses JSON data files for structured storage, and integrates with Google Sheets for cloud-based feedback collection.
- Activity QA: Ask about upcoming talks, reading groups, and events.
- Academic Offerings QA: Ask about degree programs (Bachelor, Master, PhD), application requirements, deadlines, and coordinators.
- Automated Scraping: Keeps data fresh by scraping the MCMP website.
- Rich Metadata Extraction: Automatically extracts detailed profile information including emails, office locations, hierarchical roles, and publication lists.
- Structured Data Tools (MCP): Implements an in-process Model Context Protocol (MCP) server that exposes
people.json,research.json, andraw_events.jsonas structured tools. This allows the LLM to perform precise queries (e.g., "List all events next week", "Who researches Logic?"). - Cloud Database (Feedback): User feedback is automatically saved to a Google Sheet for persistent, cloud-based storage (with a local JSON fallback).
- Multi-LLM Support: Configured to work seamlessly with Google Gemini, but also supports OpenAI and Anthropic.
- Institutional Graph: Uses a graph-based layer (
data/graph) to understand organizational structure (Chairs, Leadership) while linking people to hierarchical Research Topics. - Configurable Personality (Leopold): The chatbot's personality is defined in
prompts/personality.md, separating tone and identity from code. Edit the file to adjust behavior without touching the engine. - Agentic Workflow: Follows the
AGENTS.mdanddocs/MD_CONVENTIONS.mdprotocols for AI-assisted development.
The system answers queries through two complementary mechanisms:
- Web Scraping β JSON Data:
scripts/update_dataset.pyscrapes the MCMP website and stores structured data indata/people.json,data/research.json, anddata/raw_events.json. It also builds the Institutional Graph (data/graph/). - MCP Structured Tools: The LLM uses an in-process MCP server (
src/mcp/) to query those JSON files precisely. Tools likesearch_people,search_research, andget_eventslet the model answer structured questions (e.g., "Who is presenting next Tuesday?") without relying on fuzzy text retrieval.
-
Clone the repository:
git clone <repository-url> cd mcmp_chatbot
-
Install dependencies:
pip install -r requirements.txt
-
Configure Secrets: Create a
.streamlit/secrets.tomlfile with your API keys.For Google Gemini (Recommended):
- Get your API key from Google AI Studio.
GEMINI_API_KEY = "your-google-gemini-key"
[!NOTE] The
GEMINI_API_KEYdetermines where the LLM usage is billed. This is often a different Google Cloud project (e.g.,gen-lang-client...) than the Service Account used for Sheets. To consolidate billing, link your API key to themcmp-chatbotproject in Google AI Studio.For Cloud Feedback (Google Sheets):
- Create a project in Google Cloud Console.
- Enable the Google Sheets API and Google Drive API.
- Create a Service Account and download the JSON key.
[gcp_service_account] type = "service_account" project_id = "..." private_key = "..." client_email = "..." # ... (other standard GCP credentials) sheet_name = "MCMP Feedback"
-
Run the Application:
streamlit run app.py
To keep the chatbot up to date with the latest MCMP events and personnel, run the update protocol:
python scripts/update_dataset.pyThis script will:
- Scrape the MCMP website (Events, People, Research).
- Accumulate JSON datasets (
data/*.json) β existing entries are updated or kept; entries are never removed. - Enrich Metadata: Run internal utilities to extract structured metadata (dates, roles) from text descriptions.
- Rebuild the Institutional Graph (
data/graph/mcmp_graph.mdandmcmp_jgraph.json).
Important
Accumulation, not replacement. The datasets grow monotonically. If an event disappears from the website (e.g. dynamic "Load more" button not triggered, or the event page is taken down), the entry is still preserved in the JSON file. The scraping_logs.json "removed" field records what was absent in the current scrape but does not reflect a deletion from the dataset.
The user interface is built entirely in Streamlit, providing a clean, responsive chat interface. It handles user sessions, admin access (password protected), and feedback forms directly in the browser.
Streamlit limits raw HTML <a href="..."> links from natively triggering backend Python callbacks securely without hard page reloads. To maintain our premium presentation while preserving instant conversational injections, the application uses a Pure Native Calendar UI:
- The calendar grid is dynamically constructed using strictly native
st.columnsandst.buttoncomponents to ensure perfect layout alignment and fast, socket-driven session behavior. - We map Streamlit's built-in button types to represent different states:
type="primary"(Today),type="secondary"(Normal/Event Day), andtype="tertiary"(Empty padding to maintain grid shape). - Event days are visually indicated natively using standard Unicode emojis (
π΅) appended to the button text string, removing the need for complex and fragile DOM-breaking CSS injections. - We target these specific built-in component types using scoped CSS pseudo-selectors (like
[data-testid="column"] button) injected viast.markdown(unsafe_allow_html=True). This cleanly overrides standard padding and sizing (using!importanttags and uniform background styles) to create a perfectly square, tight, and consistent grid without breaking Streamlit's strict React DOM behavior. - Clicking a date silently injects a hidden prompt ("What events are scheduled for X?") into the chat sequence, triggering real-time MCP tool calls directly within the existing chat viewer.
The core logic (src/core/engine.py) connects to the Gemini API (or others) to generate responses. It offers the model a set of MCP tools; the model decides whether and how to call them based on the user's query.
- JSON Data Files: Scraped content is stored as structured JSON in
data/people.json,data/research.json,data/raw_events.json, anddata/academic_offerings.json. These are the source of truth queried by the MCP tools. - Institutional Graph: Organizational relationships are stored in
data/graph/mcmp_graph.mdanddata/graph/mcmp_jgraph.json, parsed bysrc/core/graph_utils.pyfor context injection. - Cloud Feedback: User feedback is pushed to Google Sheets via the Google Drive API, acting as a cloud database for ongoing user data collection.
The system connects five key data types to answer complex questions:
- People (
data/people.json): Comprehensive profiles including bios, contact details (email, phone, office), organizational roles, and selected publications. - Research (
data/research.json): Hierarchical structure of research areas (e.g., Logic, Philosophy of Science) and their subtopics, with automated linking to people. - Events (
data/raw_events.json): Upcoming talks and workshops. - Academic Offerings (
data/academic_offerings.json): Structured program info for Bachelor, Master, and PhD pathways β including ECTS, deadlines, coordinators, required documents, and contact emails. Scraped from the MCMP "For Students" section at most once per 30 days. - Institutional Graph (
data/graph/mcmp_graph.md): A knowledge graph that links People to Organizational Units (Chairs) and defines hierarchy (e.g., who leads a chair, who supervises whom).
How they interact:
- When a user asks "Who works at the Chair of Philosophy of Science?", the Graph identifies the Chair entity and its
affiliated_withedges. - The system then retrieves detailed profiles from People data via the
search_peopleMCP tool. - If the user asks "What does Ignacio Ojea research?", the
search_peopletool returns his full profile including linked Research Topics.
To handle specific queries that require structured data access (e.g., "Which events are happening between date X and Y?"), the system implements a lightweight MCP Server (src/mcp/).
- Tools: Exposes Python functions (
search_people,search_research,get_events,search_graph,search_academic_offerings) as tools to the LLM. - Execution: The engine offers these tools to the LLM. If the LLM determines it needs data, it calls the tool, and the result is fed back for the final answer.
- Toggle: This feature can be enabled/disabled via the Streamlit sidebar to manage latency and costs.
(See
docs/MCP_AGENT.mdfor full implementation details)
To ensure reliable tool usage, the system implements:
- Dynamic Injection: Tools are explicitly listed in the system prompt.
- Force Usage: Imperative commands ("Just check") prevent the LLM from asking for permission.
- Data Enrichment: If context is partial (e.g., missing abstracts), the LLM is mandated to use tools to fetch full details.
This diagram illustrates how the system combines Scraping (Data Freshness), Graph (Relationships), and MCP (Structured Queries) to answer a user query.
graph TD
UserQuery[User Query] --> LLM[LLM - Gemini]
subgraph "Data Layer (kept fresh by scraper)"
JsonDB[(data/people.json\ndata/research.json\ndata/raw_events.json)]
GraphDB[(data/graph/\nmcmp_graph.md\nmcmp_jgraph.json)]
end
subgraph "Tool Execution (MCP)"
LLM -- "Calls tool?" --> ToolCall{Decision}
ToolCall -- Yes --> ExecuteTool[Execute MCP Tool\nsearch_people / get_events\nsearch_research / search_graph]
ExecuteTool --> JsonDB
ExecuteTool --> GraphDB
JsonDB --> |Structured Result| LLM
GraphDB --> |Graph Result| LLM
ToolCall -- No/Done --> GenerateAnswer[Generate Answer]
end
GenerateAnswer --> FinalResponse[Final Response]
style JsonDB fill:#e8f5e9,stroke:#1b5e20
style GraphDB fill:#f3e5f5,stroke:#4a148c
style LLM fill:#fff3e0,stroke:#e65100
- User Query: The user asks a question in Streamlit.
- LLM Evaluation: The LLM receives the query and a list of available MCP tools.
- Tool Calls (if needed):
- Scenario A (Simple Query): The LLM answers from its general knowledge + graph context injected in the system prompt.
- Scenario B (Needs Structured Data): The LLM calls a tool (e.g.,
get_events(query="next week")). The system executes this against the JSON Data Files and feeds the precise result back.
- Final Answer: The LLM synthesizes tool outputs and graph context into the final response.
Query: "What is Hannes Leitgeb working on and what are his upcoming events?"
-
LLM receives query with the available MCP tools listed.
-
Tool Call 1 β
search_people:- The LLM calls:
search_people(query="Hannes Leitgeb") - Result: Full profile β research areas, publications, office details.
- The LLM calls:
-
Tool Call 2 β
get_events:- The LLM calls:
get_events(query="Hannes Leitgeb") - Result:
[{"title": "Talk at LMU", "date": "2024-10-15", ...}]
- The LLM calls:
-
Final Synthesis:
"Hannes Leitgeb is currently working on Logic and Probability (from
search_people). He leads the Chair of Logic and Philosophy of Language (from Graph). Regarding his schedule, he has an upcoming talk at LMU on October 15th (fromget_events)."
mcmp_chatbot/
βββ app.py # Main Streamlit application entry point
βββ src/
β βββ core/ # AI engine (Gemini), Graph utils, Personality loader
β βββ mcp/ # MCP Tools and Server implementation
β βββ scrapers/ # Scrapers for MCMP website
β βββ ui/ # Streamlit UI components
β βββ utils/ # Helper functions (logging, etc.)
βββ prompts/ # Chatbot personality configuration
β βββ personality.md # Leopold's identity, tone, and guidelines
βββ data/ # Local data storage (JSONs, Graph)
β βββ people.json
β βββ research.json
β βββ raw_events.json
β βββ graph/ # Institutional graph (md + json)
βββ docs/ # Project documentation and proposals
β βββ MCP_AGENT.md
β βββ HOUSEKEEPING.md # Maintenance protocols
β βββ PERSONALITY_AGENT.md
β βββ MD_CONVENTIONS.md # Markdown conventions
βββ scripts/ # Maintenance and update scripts
β βββ update_dataset.py # Main data update script
βββ tests/ # Unit and integration tests
βββ AGENTS.md # Guidelines for AI Agents
βββ requirements.txt # Python dependencies
This project uses a structured workflow for AI agents.
- AGENTS.md: Read this first if you are an AI assistant.
- docs/MD_CONVENTIONS.md: Defines the schema for Markdown files and task management.