Turns PDF books into audiobooks — and more. Upload PDFs, pick a voice, and get chapter-marked MP3s, AI digests, translations, PDF/EPUB exports, and read-along synced EPUBs (audio + highlighted text) you can listen to offline on a phone.
Built for local use on Apple Silicon Macs. Fully offline after the initial model downloads (AI features need a DeepSeek API key).
- PDF → audiobook: chapter detection (deterministic tiers + optional LLM TOC detection), per-chapter TTS synthesis, single MP3 assembly with ID3v2 chapter markers.
- Raw-first uploads: every upload gets instant
pdftotextraw text; the slow Marker extraction (OCR-capable) is opt-in and can run later. - Per-chapter control: edit text, re-synthesize, include/exclude, suspend/queue, AI cleanup of OCR artifacts, manual or LLM-proposed chapter boundaries.
- Translations: first-class per-language chapter variants (DeepSeek) with their own TTS audio and assemblies; the original text is always preserved. Translation streams live into the side-by-side view — you see the model's thinking, then the translated text token by token.
- Ask AI + notes: whole-book or per-chapter prompts; every answer is auto-saved as a note on the book, and any note can be appended to the book as a chapter of its own — ready to reorder and synthesize.
- Digest books: select N books → one synthetic book with an AI summary chapter per source, ready to synthesize.
- External API: plain JSON endpoints (
POST /api/books, seedocs/synthetic-books-api.md) so scripts and other projects can create synthetic books and chapters — with optional straight-to-audio synthesis. Ships withscripts/hn-top10.mjs, which turns any day's top Hacker News stories (via hckrnews.com archives) into a podcast-style book — one chapter per story, article text extracted with Defuddle, community reaction capped at 20%. - Document export: selected chapters as PDF/EPUB (Vivliostyle), or as a synced EPUB — EPUB 3 with Media Overlays: embedded audio plus sentence-level highlighted text, valid per epubcheck.
- Read-along on iPhone: a self-hosted Storyteller companion (see
storyteller/) auto-imports synced EPUBs; the free Storyteller Reader app downloads them for fully offline listening with live text highlighting. - Library organization: nested folders with drag & drop, cross-folder search, lightweight profiles (workspaces) so different people keep separate libraries.
- Library chat: an agentic assistant (
/chat) that searches the content of every book — hybrid full-text + semantic search (local BGE-M3 embeddings, cross-language: ask in English, find the Bulgarian passage and vice versa) — and streams answers with verified citations. Click a source chip to open the PDF at that page, the chapter, or the translation view. Answers can be saved as notes.
Upload → rawExtract (pdftotext, seconds, always)
→ extract (Marker, opt-in, OCR-capable) → normalize → synthesize (TTS) → assemble → MP3
→ translate → synthesizeTranslation → per-language assembly
→ assembleDocument → PDF / EPUB / synced EPUB
Jobs run through Graphile Worker in six pools (TTS, raw text, extraction, assembly, AI/translation, search indexing) with maxAttempts: 1 — nothing retries silently; the user reviews failures and decides. Chapter text falls back customText ?? cleanText ?? rawText at synthesis time.
TTS engines: Kokoro (English + 8 more languages), KugelAudio (24 EU languages incl. Bulgarian, local 4-bit MLX quant), BG-TTS V5 MLX, and Meta MMS Bulgarian — all local, GPU-accelerated via MPS/Metal.
During synthesis the server keeps a per-chunk text↔audio timing map (chNNN.sync.json) next to each MP3. That map powers the web UI's read-along player and the synced EPUB export — and once it exists, the intermediate chunk WAVs can be deleted to reclaim disk.
pnpm monorepo: packages/server (Fastify + tRPC + Graphile Worker + Drizzle/Postgres, port 3034) and packages/web (React 19 + Vite + Tailwind v4 + react-router 7, port 3033). Python TTS/extraction scripts live in scripts/; the optional Storyteller companion in storyteller/.
The detailed, maintained map of files, tables, routes, and pipeline internals is in AGENTS.md — this README stays intentionally high-level.
PostgreSQL 17 with pgvector in Docker (pgvector/pgvector:pg17, host port 5433), schema via Drizzle ORM: profiles, folders, books, book_files, chapters, chapter_translations, assemblies, documents, notes, book_logs, book_chunks (search index: FTS + embeddings). See AGENTS.md for column-level docs. Migrations: pnpm db:generate + pnpm db:migrate.
The compose file switched from postgres:17-alpine to pgvector/pgvector:pg17 (needed for the vector extension). The images use different C libraries (musl vs glibc), so the data volume is not reused — the compose volume was renamed pgdata → pgdata17 and data moves via dump/restore:
# 1. While still on the old container:
docker compose exec postgres pg_dump -U pdf2audio --no-owner pdf2audio > pgbackup.sql
# 2. Pull up the new image + fresh volume (compose file already updated):
docker compose up -d
# 3. Restore:
docker compose exec -T postgres psql -U pdf2audio -d pdf2audio < pgbackup.sql
# 4. Verify the app, then eventually: docker volume rm pdf2audio_pgdataThe old pdf2audio_pgdata volume stays untouched as a rollback until you delete it.
After the restore, index the library for search: pnpm db:migrate && pnpm backfill:index (FTS is available within minutes; BGE-M3 embeddings fill in as a background pass).
All runtime data lives in ./data/ (gitignored, resolved relative to packages/server):
data/uploads/{bookId}/ Uploaded PDFs
data/tmp/{bookId}/ Marker JSON output
data/output/{bookId}/ Chapter MP3s + sync maps, assemblies, exported documents
data/output/{bookId}/{lang}/ Translation audio
data/output/{bookId}/chunks/ Chunk WAV previews (disposable once sync maps exist)
data/previews/ Voice preview MP3s
- Node.js >= 20 and pnpm
- Python 3.10+ with a conda environment (or global pip)
- Docker (for Postgres and optionally Storyteller)
- FFmpeg —
brew install ffmpeg - poppler (
pdftotext) —brew install poppler - espeak-ng —
brew install espeak-ng - Marker —
pip install marker-pdf==1.8.5 - Kokoro —
pip install kokoro soundfile - Bulgarian narrator —
pip install mlx numpy huggingface_hubandpip install "nanocodec-mlx @ git+https://github.com/nineninesix-ai/nanocodec-mlx.git" - Meta MMS Bulgarian —
pip install transformers torch - KugelAudio narrator —
pip install mlx-audio, thenpip install "transformers==4.57.6" "regex<2025.0.0"(mlx-audio pulls transformers 5.x, which breaks marker-pdf) - Optional: a DeepSeek API key for translation, cleanup, digests, Ask AI, and LLM chapter detection
# Clone and install everything
pnpm setup
# Copy env file (defaults work out of the box; add DEEPSEEK_API_KEY for AI features)
cp .env.example .env
# Start Postgres
pnpm db:up
# Run migrations
pnpm db:migrate
# Start dev servers (server on :3034, web on :3033)
pnpm devcd storyteller
openssl rand -base64 32 > STORYTELLER_SECRET_KEY.txt
docker compose up -d # web UI + API on http://localhost:8001Create the admin account at http://localhost:8001, then set READALOUD_DROP_DIR=<repo>/storyteller/data/import in .env — the "Copy to Storyteller import folder" checkbox on synced-EPUB exports will drop files there and Storyteller auto-imports them. Install the free Storyteller Reader iOS/Android app and point it at your Mac's LAN address on port 8001.
pnpm dev # Start server + web in parallel
pnpm dev:server # Server only (port 3034)
pnpm dev:web # Web only (port 3033)
pnpm db:up # Start Postgres in Docker
pnpm db:down # Stop Postgres
pnpm db:generate # Generate Drizzle migration from schema changes
pnpm db:migrate # Apply migrations
pnpm setup # Full setup (system deps check, Python/Node deps, data dirs)
pnpm jobs # Show Graphile Worker queue status
pnpm jobs:clear # Delete all queued jobs
cd packages/server && pnpm test # Server test suite (spins up template DB, runs migrations)- Docker Postgres is mapped to host port 5433 to avoid conflicts with other Postgres instances on 5432.
- The Kokoro model (
hexgrad/Kokoro-82M, 82M params, Apache-2.0) auto-downloads on first run;HF_HUB_OFFLINE=1is set afterwards, so models must be cached before offline use. - The Bulgarian-capable narrators are
BG-TTS V5 (Radi Totev MLX port),MMS Bulgarian (Meta), andKugelAudio (7B, 24 EU languages); Bulgarian voice speed is fixed (UI disables the slider). - KugelAudio (
kugelaudio/kugelaudio-0-open, Apache-2.0) runs from a local 4-bit MLX quantization (~5 GB) at~/.cache/pdf2audio-models/kugelaudio-0-open-4bit(override withKUGEL_TTS_MODEL_PATH);pnpm setupdownloads and converts it. ~1.5x realtime on an M4 Pro. facebook/mms-tts-bulis licensedCC-BY-NC-4.0.- Best Kokoro voices:
af_heart(A tier),af_bella(A- tier),bf_emma(B- tier). - Synced EPUBs deliberately end with a non-narrated colophon page — it works around a crash in the Storyteller iOS app when the last spine item carries a media overlay (reported upstream).