A secure REST API that answers natural-language questions from your own documents, with cited sources. Upload PDFs or text files (policies, manuals, notes); DocMind indexes them and answers questions like "How many vacation days do I get?" — quoting the exact file and page the answer came from.
Built with Java 17, Spring Boot 4, Spring AI, PostgreSQL, and JWT security. This is a working implementation of RAG (Retrieval-Augmented Generation) — the pattern behind enterprise document assistants.
flowchart LR
subgraph Ingestion["Ingestion (on upload)"]
A[PDF / TXT / MD] --> B[Extract text per page<br/>PDFBox]
B --> C[Split into overlapping<br/>~1000-char chunks]
C --> D[Embed each chunk<br/>OpenAI text-embedding-3-small]
D --> E[(PostgreSQL<br/>chunks + vectors)]
end
subgraph Question["Question (on /api/ask)"]
Q[Question] --> QE[Embed question]
QE --> R[Cosine similarity vs<br/>all owned chunks → top 5]
E --> R
R --> P[Prompt: answer ONLY<br/>from these excerpts]
P --> L[Chat model<br/>gpt-4o-mini]
L --> ANS[Answer + citations<br/>file, page, snippet, score]
end
In plain English: on upload, the librarian tears your document into paragraph-sized index cards and files each card by meaning. When you ask a question, it pulls the 5 closest cards and hands them to a writer with strict orders: answer only from these cards, or say you don't know. The AI never free-styles from its own memory — that's what prevents hallucinated answers.
| Method | Endpoint | Auth | Purpose |
|---|---|---|---|
| POST | /api/auth/register |
— | Create account → JWT |
| POST | /api/auth/login |
— | Log in → JWT |
| POST | /api/documents |
Bearer | Upload .pdf/.txt/.md (multipart file), chunked + embedded immediately |
| GET | /api/documents |
Bearer | List your documents |
| DELETE | /api/documents/{id} |
Bearer | Delete a document and its index |
| POST | /api/ask |
Bearer | {"question": "..."} → answer + citations |
| GET | /api/health |
— | Liveness check |
Interactive docs: http://localhost:8080/swagger-ui.html (use the Authorize button with your JWT).
Example /api/ask response:
{
"answer": "Employees receive 25 vacation days per year [1].",
"citations": [
{
"filename": "employee-handbook.pdf",
"page": 12,
"snippet": "All full-time employees receive 25 vacation days per calendar year...",
"score": 0.874
}
]
}Prerequisites: Java 17, PostgreSQL running locally, an OpenAI API key.
# 1. Create the database (once)
psql -U postgres -c "CREATE DATABASE docmind"
# 2. Set your API key for this shell session
$env:OPENAI_API_KEY = "sk-..."
# 3. Run
.\mvnw.cmd spring-boot:runDefaults assume Postgres at localhost:5432 with postgres/postgres; override via DB_URL, DB_USERNAME, DB_PASSWORD env vars. The app boots without an API key, but AI endpoints return 502 until one is set.
.\mvnw.cmd test27 tests, no network and no API key needed: the AI calls sit behind two small interfaces (AiChat, Embedder) that tests replace with mocks, and the database is in-memory H2.
- AI behind a seam — services depend on two tiny interfaces, not on Spring AI directly. Tests mock them; swapping OpenAI for Claude or a local model touches one adapter class.
- Citations by construction — every chunk keeps its source file and page through the whole pipeline, so answers can always show receipts.
- Per-user isolation — retrieval only ever searches documents the authenticated user owns.
- Honest failure modes — wrong file type → 415, unreadable PDF → 422, oversized upload → 413, AI provider down → 502 with a clean JSON body. Never a stack trace.
See DECISIONS.md for the trade-offs behind each choice (chunk size, in-Java cosine similarity vs pgvector, sync vs async ingestion).
Java 17 · Spring Boot 4.1 · Spring AI 2.0 (OpenAI chat + embeddings) · Spring Security (stateless JWT) · Spring Data JPA · PostgreSQL · PDFBox · springdoc-openapi · JUnit 5 + Mockito + AssertJ · H2 (tests)
- pgvector for SQL-native similarity search at scale
- Async ingestion with upload status polling for large documents
- Conversation memory (follow-up questions)
- Rate limiting per user