Skip to content

Repository files navigation

Memex

Python FastAPI Qdrant Neo4j Whisper CLIP Next.js Docker Ollama License: MIT

A fully local, multimodal RAG system that indexes your files — PDFs, audio, video, images — and lets you chat with them, search visually, and explore connections between ideas. Nothing leaves your machine.


Demo

Memex demo — ingest and query walkthrough


Why I built this

I accumulated years of research papers, meeting notes, and articles across dozens of folders — retrieving the right context when I needed it took longer than the thinking itself. I wanted a system that ran entirely on my own hardware with no cloud subscriptions, no data leaving my machine, and no per-query cost. Memex is the result: drop a file in a folder and within seconds it is chunked, embedded, and ready to chat with via a streaming interface that cites its sources.


What it does

Frontend Port Purpose
Second Brain 3001 Chat with your documents. Streaming answers, citation badges, entity graph, query history.
Media Intel 3000 Visual search over images and video. Drag-drop ingest, scene detection, text-to-image search.
Backend API 8000 All ingest and query endpoints. Interactive docs at /docs.
Qdrant 6333 Vector database dashboard.
Neo4j 7474 Knowledge graph browser.

Prerequisites

Requirement Notes
Docker Desktop Enable WSL 2 backend on Windows
Ollama Runs on the host, not inside Docker
NVIDIA GPU (recommended) Whisper large-v3 and CLIP ViT-L/14 are GPU-accelerated. CPU fallback works but is slow.
NVIDIA Container Toolkit Required for GPU pass-through into the worker container

Pull required Ollama models once:

ollama pull llama3.2:latest
ollama pull nomic-embed-text

Quick start

# 1. Clone (or open an existing clone)
cd my-second-brain

# 2. Start all services
docker compose -f docker/docker-compose.yml up -d

# 3. Start the Second Brain frontend (new terminal)
cd frontend-second-brain
npm install   # first time only
npm run dev

# 4. Start the Media Intel frontend (new terminal)
cd frontend-media-intel
npm install   # first time only
npm run dev

Open:

The backend and worker start automatically via Docker. The file watcher starts with the worker container and watches the folders defined in .env.


Watch folders

The worker monitors two folders on your machine for new files:

Folder Destination
%USERPROFILE%\SecondBrain\ second_brain_text + second_brain_audio collections
%USERPROFILE%\MediaIntel\ media_images + media_video_keyframes collections

Drop any supported file into either folder and it is ingested automatically within a few seconds.

Supported formats: PDF, DOCX, PPTX, TXT, MD, HTML, MP3, WAV, M4A, OGG, FLAC, MP4, MOV, MKV, AVI, WEBM, JPG, PNG, GIF, WEBP, TIFF, BMP


Ingesting files manually

Via the frontend: Use the Ingest tab (📥) in Second Brain or the drag-drop zone in Media Intel.

Via the API:

# Upload a single file
curl -X POST http://localhost:8000/ingest/upload -F "file=@C:\path\to\notes.pdf"

# Queue a folder path (files already on the server's watch volume)
curl -X POST http://localhost:8000/ingest/folder `
  -H "Content-Type: application/json" `
  -d '{"folder_path": "/watched_home/SecondBrain/research"}'

# Check task status
curl http://localhost:8000/ingest/status/{task_id}

Querying

Chat (streaming):

curl -X POST http://localhost:8000/query/chat/stream `
  -H "Content-Type: application/json" `
  -d '{"query": "What are the main arguments in my notes on consciousness?", "collection": "second_brain_text", "top_k": 5}'

Visual search:

# Text → images
curl -X POST http://localhost:8000/query/visual `
  -H "Content-Type: application/json" `
  -d '{"query": "sunset over mountains", "top_k": 10}'

# Image → similar images
curl -X POST http://localhost:8000/query/visual/image -F "file=@photo.jpg"

# Video scene search
curl -X POST http://localhost:8000/query/video `
  -H "Content-Type: application/json" `
  -d '{"query": "person walking on beach"}'

Configuration

All settings live in .env at the project root. Key variables:

# Models
LLM_MODEL=llama3.2:latest          # Ollama chat model
EMBED_MODEL=nomic-embed-text        # Ollama embedding model
WHISPER_MODEL=large-v3             # faster-whisper model size
CLIP_MODEL=ViT-L-14                # CLIP model for image/video

# Retrieval
RETRIEVAL_TOP_K=20                 # candidates retrieved by hybrid search
RERANK_TOP_K=5                     # top-N kept after cross-encoder re-ranking

# Hardware
WHISPER_DEVICE=cuda                # cuda or cpu
WHISPER_COMPUTE_TYPE=float16       # float16 (GPU) or int8 (CPU)

# Remote access
TAILSCALE_HOSTNAME=                # e.g. your-pc.tail12345.ts.net — see docs/TAILSCALE.md

Pipeline overview

File dropped
    │
    ▼
Celery worker
    ├── PDF/DOCX/TXT  → TextExtractor   → 512-token chunks → nomic-embed (dense) + BM25 (sparse)
    ├── MP3/WAV/M4A   → AudioExtractor  → Whisper large-v3 → 30s chunks with timestamps
    ├── JPG/PNG/…     → ImageExtractor  → CLIP ViT-L/14 + BLIP caption + EXIF
    └── MP4/MOV/…     → VideoExtractor  → scene cuts → keyframe CLIP + aligned transcript
         │
         ▼
    Qdrant (4 collections)    Neo4j (entity graph)    SQLite (file hashes)
         │
         ▼
Query: hybrid search (RRF) → cross-encoder re-rank → Ollama llama3.2 → streamed answer

Architecture

graph LR
    SB["Second Brain :3001"]
    MI["Media Intel :3000"]
    API["FastAPI :8000"]
    W["Celery Worker"]
    Q["Qdrant :6333"]
    N["Neo4j :7474"]
    R["Redis :6379"]

    SB -->|HTTP| API
    MI -->|HTTP| API
    API -->|vector search| Q
    API -->|graph query| N
    API -->|enqueue task| R
    R -->|dequeue| W
    W -->|upsert vectors| Q
    W -->|write entities| N
    W -->|store result| R
Loading

Collections

Collection Content
second_brain_text PDF, DOCX, PPTX, TXT, MD, HTML chunks
second_brain_audio Audio transcript chunks with timestamps
media_images Image CLIP embeddings + captions
media_video_keyframes Video keyframe CLIP embeddings + aligned transcript

Maintenance

Apply INT8 quantization to existing collections (run once after the first ingest, or after upgrading from an older version):

docker exec backend python scripts/apply_quantization.py --host qdrant

This compresses all vector data to INT8 in-place — no re-ingestion needed. RAM usage drops ~4× over the next few minutes as Qdrant rewrites its segments.

View logs:

docker compose -f docker/docker-compose.yml logs -f backend
docker compose -f docker/docker-compose.yml logs -f worker

Restart a single service:

docker compose -f docker/docker-compose.yml restart backend

Eval harness

Measures retrieval recall@5 against a set of known queries:

# 1. Edit scripts/eval_queries.yaml to add your (query, expected_source) pairs
# 2. Run inside the backend container
docker exec backend python scripts/eval_retrieval.py

Prints a per-query hit/miss table and overall recall@5 score. Target: ≥ 80%.


Remote access via Tailscale

Access Second Brain from your phone or laptop without port forwarding. See docs/TAILSCALE.md for full setup instructions.

Short version:

  1. Install Tailscale on this machine and sign in
  2. Set TAILSCALE_HOSTNAME=your-pc.tail12345.ts.net in .env
  3. Create .env.local in each frontend pointing NEXT_PUBLIC_API_URL at the Tailscale hostname
  4. Install Tailscale on your phone and open http://your-pc.tail12345.ts.net:3001

Project structure

my-second-brain/
├── backend/                  FastAPI app + Celery workers
│   ├── api/                  Route handlers (chat, ingest, query, stats, history, graph)
│   ├── extractors/           TextExtractor, AudioExtractor, ImageExtractor, VideoExtractor
│   ├── retrievers/           hybrid_search.py, reranker.py
│   ├── graph/                GLiNER entity extraction, Neo4j writer + retriever
│   ├── db/                   hash_store.py (SQLite)
│   └── workers/              celery_app.py, file_watcher.py
├── docker/                   docker-compose.yml + DB storage volumes
├── frontend-second-brain/    Next.js 16, port 3001
├── frontend-media-intel/     Next.js 16, port 3000
├── scripts/                  apply_quantization.py, eval_retrieval.py, test_ui.html
├── docs/                     ARCHITECTURE.md, PHASES.md, TAILSCALE.md, TECH_STACK.md
└── .env                      All configuration

Troubleshooting

Backend won't start — Neo4j not healthy Neo4j takes up to 40 seconds on first boot. Run docker compose logs neo4j and wait for Started before retrying.

Worker OOM during Whisper transcription The worker container is capped at 10 GB RAM. If it crashes, reduce WHISPER_MODEL to medium in .env and restart.

Chat returns no answer / blank response Check docker compose logs backend. If you see CUDA out of memory, the LLM ran out of VRAM. Switch to a smaller model: set LLM_MODEL=llama3.2:latest (already the default) or LLM_MODEL=qwen2.5:3b.

nomic-embed-text not found Ollama runs on the host. Run ollama pull nomic-embed-text in a host terminal, not inside Docker.

Qdrant dashboard shows empty collections after restart Normal — collections are created at startup by qdrant_init.py. If points are missing, check docker compose logs worker for ingest errors.

About

Fully local multimodal RAG system — chat with PDFs, audio, video and images. No cloud, no API keys.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages