A fully local, multimodal RAG system that indexes your files — PDFs, audio, video, images — and lets you chat with them, search visually, and explore connections between ideas. Nothing leaves your machine.
I accumulated years of research papers, meeting notes, and articles across dozens of folders — retrieving the right context when I needed it took longer than the thinking itself. I wanted a system that ran entirely on my own hardware with no cloud subscriptions, no data leaving my machine, and no per-query cost. Memex is the result: drop a file in a folder and within seconds it is chunked, embedded, and ready to chat with via a streaming interface that cites its sources.
| Frontend | Port | Purpose |
|---|---|---|
| Second Brain | 3001 | Chat with your documents. Streaming answers, citation badges, entity graph, query history. |
| Media Intel | 3000 | Visual search over images and video. Drag-drop ingest, scene detection, text-to-image search. |
| Backend API | 8000 | All ingest and query endpoints. Interactive docs at /docs. |
| Qdrant | 6333 | Vector database dashboard. |
| Neo4j | 7474 | Knowledge graph browser. |
| Requirement | Notes |
|---|---|
| Docker Desktop | Enable WSL 2 backend on Windows |
| Ollama | Runs on the host, not inside Docker |
| NVIDIA GPU (recommended) | Whisper large-v3 and CLIP ViT-L/14 are GPU-accelerated. CPU fallback works but is slow. |
| NVIDIA Container Toolkit | Required for GPU pass-through into the worker container |
Pull required Ollama models once:
ollama pull llama3.2:latest
ollama pull nomic-embed-text# 1. Clone (or open an existing clone)
cd my-second-brain
# 2. Start all services
docker compose -f docker/docker-compose.yml up -d
# 3. Start the Second Brain frontend (new terminal)
cd frontend-second-brain
npm install # first time only
npm run dev
# 4. Start the Media Intel frontend (new terminal)
cd frontend-media-intel
npm install # first time only
npm run devOpen:
- Second Brain → http://localhost:3001
- Media Intel → http://localhost:3000
- API docs → http://localhost:8000/docs
The backend and worker start automatically via Docker. The file watcher starts with the worker container and watches the folders defined in .env.
The worker monitors two folders on your machine for new files:
| Folder | Destination |
|---|---|
%USERPROFILE%\SecondBrain\ |
second_brain_text + second_brain_audio collections |
%USERPROFILE%\MediaIntel\ |
media_images + media_video_keyframes collections |
Drop any supported file into either folder and it is ingested automatically within a few seconds.
Supported formats: PDF, DOCX, PPTX, TXT, MD, HTML, MP3, WAV, M4A, OGG, FLAC, MP4, MOV, MKV, AVI, WEBM, JPG, PNG, GIF, WEBP, TIFF, BMP
Via the frontend: Use the Ingest tab (📥) in Second Brain or the drag-drop zone in Media Intel.
Via the API:
# Upload a single file
curl -X POST http://localhost:8000/ingest/upload -F "file=@C:\path\to\notes.pdf"
# Queue a folder path (files already on the server's watch volume)
curl -X POST http://localhost:8000/ingest/folder `
-H "Content-Type: application/json" `
-d '{"folder_path": "/watched_home/SecondBrain/research"}'
# Check task status
curl http://localhost:8000/ingest/status/{task_id}Chat (streaming):
curl -X POST http://localhost:8000/query/chat/stream `
-H "Content-Type: application/json" `
-d '{"query": "What are the main arguments in my notes on consciousness?", "collection": "second_brain_text", "top_k": 5}'Visual search:
# Text → images
curl -X POST http://localhost:8000/query/visual `
-H "Content-Type: application/json" `
-d '{"query": "sunset over mountains", "top_k": 10}'
# Image → similar images
curl -X POST http://localhost:8000/query/visual/image -F "file=@photo.jpg"
# Video scene search
curl -X POST http://localhost:8000/query/video `
-H "Content-Type: application/json" `
-d '{"query": "person walking on beach"}'All settings live in .env at the project root. Key variables:
# Models
LLM_MODEL=llama3.2:latest # Ollama chat model
EMBED_MODEL=nomic-embed-text # Ollama embedding model
WHISPER_MODEL=large-v3 # faster-whisper model size
CLIP_MODEL=ViT-L-14 # CLIP model for image/video
# Retrieval
RETRIEVAL_TOP_K=20 # candidates retrieved by hybrid search
RERANK_TOP_K=5 # top-N kept after cross-encoder re-ranking
# Hardware
WHISPER_DEVICE=cuda # cuda or cpu
WHISPER_COMPUTE_TYPE=float16 # float16 (GPU) or int8 (CPU)
# Remote access
TAILSCALE_HOSTNAME= # e.g. your-pc.tail12345.ts.net — see docs/TAILSCALE.mdFile dropped
│
▼
Celery worker
├── PDF/DOCX/TXT → TextExtractor → 512-token chunks → nomic-embed (dense) + BM25 (sparse)
├── MP3/WAV/M4A → AudioExtractor → Whisper large-v3 → 30s chunks with timestamps
├── JPG/PNG/… → ImageExtractor → CLIP ViT-L/14 + BLIP caption + EXIF
└── MP4/MOV/… → VideoExtractor → scene cuts → keyframe CLIP + aligned transcript
│
▼
Qdrant (4 collections) Neo4j (entity graph) SQLite (file hashes)
│
▼
Query: hybrid search (RRF) → cross-encoder re-rank → Ollama llama3.2 → streamed answer
graph LR
SB["Second Brain :3001"]
MI["Media Intel :3000"]
API["FastAPI :8000"]
W["Celery Worker"]
Q["Qdrant :6333"]
N["Neo4j :7474"]
R["Redis :6379"]
SB -->|HTTP| API
MI -->|HTTP| API
API -->|vector search| Q
API -->|graph query| N
API -->|enqueue task| R
R -->|dequeue| W
W -->|upsert vectors| Q
W -->|write entities| N
W -->|store result| R
| Collection | Content |
|---|---|
second_brain_text |
PDF, DOCX, PPTX, TXT, MD, HTML chunks |
second_brain_audio |
Audio transcript chunks with timestamps |
media_images |
Image CLIP embeddings + captions |
media_video_keyframes |
Video keyframe CLIP embeddings + aligned transcript |
Apply INT8 quantization to existing collections (run once after the first ingest, or after upgrading from an older version):
docker exec backend python scripts/apply_quantization.py --host qdrantThis compresses all vector data to INT8 in-place — no re-ingestion needed. RAM usage drops ~4× over the next few minutes as Qdrant rewrites its segments.
View logs:
docker compose -f docker/docker-compose.yml logs -f backend
docker compose -f docker/docker-compose.yml logs -f workerRestart a single service:
docker compose -f docker/docker-compose.yml restart backendMeasures retrieval recall@5 against a set of known queries:
# 1. Edit scripts/eval_queries.yaml to add your (query, expected_source) pairs
# 2. Run inside the backend container
docker exec backend python scripts/eval_retrieval.pyPrints a per-query hit/miss table and overall recall@5 score. Target: ≥ 80%.
Access Second Brain from your phone or laptop without port forwarding. See docs/TAILSCALE.md for full setup instructions.
Short version:
- Install Tailscale on this machine and sign in
- Set
TAILSCALE_HOSTNAME=your-pc.tail12345.ts.netin.env - Create
.env.localin each frontend pointingNEXT_PUBLIC_API_URLat the Tailscale hostname - Install Tailscale on your phone and open
http://your-pc.tail12345.ts.net:3001
my-second-brain/
├── backend/ FastAPI app + Celery workers
│ ├── api/ Route handlers (chat, ingest, query, stats, history, graph)
│ ├── extractors/ TextExtractor, AudioExtractor, ImageExtractor, VideoExtractor
│ ├── retrievers/ hybrid_search.py, reranker.py
│ ├── graph/ GLiNER entity extraction, Neo4j writer + retriever
│ ├── db/ hash_store.py (SQLite)
│ └── workers/ celery_app.py, file_watcher.py
├── docker/ docker-compose.yml + DB storage volumes
├── frontend-second-brain/ Next.js 16, port 3001
├── frontend-media-intel/ Next.js 16, port 3000
├── scripts/ apply_quantization.py, eval_retrieval.py, test_ui.html
├── docs/ ARCHITECTURE.md, PHASES.md, TAILSCALE.md, TECH_STACK.md
└── .env All configuration
Backend won't start — Neo4j not healthy
Neo4j takes up to 40 seconds on first boot. Run docker compose logs neo4j and wait for Started before retrying.
Worker OOM during Whisper transcription
The worker container is capped at 10 GB RAM. If it crashes, reduce WHISPER_MODEL to medium in .env and restart.
Chat returns no answer / blank response
Check docker compose logs backend. If you see CUDA out of memory, the LLM ran out of VRAM. Switch to a smaller model: set LLM_MODEL=llama3.2:latest (already the default) or LLM_MODEL=qwen2.5:3b.
nomic-embed-text not found
Ollama runs on the host. Run ollama pull nomic-embed-text in a host terminal, not inside Docker.
Qdrant dashboard shows empty collections after restart
Normal — collections are created at startup by qdrant_init.py. If points are missing, check docker compose logs worker for ingest errors.
