Skip to content

Repository files navigation

qwacback - Questions Worth Asking Continuously

AI-Assisted

A question bank for civil society surveys. Import DDI-Codebook XML, browse and search questions, export as DDI or XLSForm.

How It Works

The core concept is a question β€” one thing you ask a respondent. Under the hood, questions are stored as DDI Codebook 2.5 variables and variable groups (see DDI_MARKUP_GUIDE.md), but the API presents them as questions:

Question type DDI storage Example
Simple (integer, text, single choice) 1 <var> "How old are you?"
Multiple choice <varGrp type="multipleResp"> + binary <var> per option "Which devices do you own?"
Grid / Likert <varGrp type="grid"> + <var> per item "Rate your trust in: Parliament, Police, ..."
Semi-open (with "other") <varGrp type="other"> + member vars + _other text var "What is your gender?" (with free text option)

Questions belong to studies β€” a study is a survey or questionnaire with metadata (title, abstract, keywords, time period, etc.).

Data flow

DDI-XML file
  β”‚
  β”œβ”€β–Ί POST /api/validate     β†’  XSD + Schematron validation only
  β”‚
  └─► POST /api/import       β†’  Validate, parse, store in DB
                                    studies / variable_groups / variables
  β”‚
  β–Ό
GET /api/questions            β†’  Browse all questions (assembled from vars + groups)
GET /api/search/questions     β†’  Search by question text, concept, name
GET /api/studies/{id}/export  β†’  The codebook as imported
GET /api/studies/{id}/xlsform β†’  Convert to XLSForm JSON

Architecture

Three Docker services:

  • qwacback (Go/PocketBase) β€” API, database, embedded NATS server
  • ddi-emitter (Node) β€” XLSForm β†’ DDI proxy over @correlaid/formtransform
  • schematron-worker (Java/Saxon HE) β€” XSD + Schematron validation over NATS

If NATS_PORT is not set, qwacback runs without validation (import-only mode).

API

Questions

  • GET /api/questions β€” List all questions across all studies.
  • GET /api/questions/{id} β€” Get a single question with full detail: embedded study, group, and variable data (categories, prequestion text, hint, interviewer instructions, skip logic as prose in universe, etc.). No additional API calls needed for a detail view.
  • GET /api/questions/{id}/xml β€” DDI-XML fragment for a single question, cut unchanged from the stored codebook.
  • GET /api/questions/{id}/xlsform β€” XLSForm JSON for a single question.
  • GET /api/studies/{id}/questions β€” List all questions for a single study.
  • GET /api/search/questions?q=<terms> β€” Search questions by question text, concept, name, and answer type. Supports &page= and &perPage= (default 20, max 100).
    • q may hold several terms, separated by spaces or commas; a question matches if any term matches, and questions matching more terms rank first (then question text > concept > name > answer type).
    • Matching ignores case and umlaut spelling (Qualitaet = QualitΓ€t), finds word forms via German and English stemming (Zufriedenheit finds "zufrieden", Weiterempfehlung finds "weiterempfehlen"), finds terms of 5+ letters inside compounds (Vertrauen finds "Institutionenvertrauen"). Tags are searched too: <concept> elements after the first one on a <var>/<varGrp>, e.g. <concept xml:lang="de">Vertrauen</concept> on an English item, make it findable in the other language without changing its wording.
    • &study_id=<id> (repeatable) limits the search to those studies; &exclude_study=<id> (repeatable) leaves studies out, e.g. the demographic standards. The MCP tool search_questions takes the same filters.

Questions of a multilingual study also carry language (the base language of question_text) and translations (the question text in the other languages, {"en": "…"}). /api/questions/{id} also returns the full translations of each variable and of the group. The search matches translated question text like the question text.

Studies

  • GET /api/search/studies?q=<term> β€” Search studies by title, keywords, and abstract. Optional &topic=<classification> filter. Supports pagination.

Import & Validation

  • POST /api/validate β€” Validate a DDI XML file (XSD + Schematron) without importing. Body: multipart/form-data with file field.
  • POST /api/import β€” Validate and import a DDI XML file. Same body format. Requires superuser auth. Responds {"valid": true, "imported": true, "study_id": "…", "message": …} on success; imported: false when the file is valid but couldn't be stored. Imports aren't persisted across deploys; see Data lives in seed_data/.

Export & Conversion

  • GET /api/studies/{id}/export β€” Export study as validated DDI-XML download.
  • GET /api/studies/{id}/xlsform β€” Export study as XLSForm JSON.
  • POST /api/convert/ddi-to-xlsform β€” Convert a DDI XML fragment to XLSForm JSON.
  • POST /api/convert/xlsform-to-ddi β€” Convert XLSForm JSON to DDI XML.

For conversion details, see CONVERSION_API.md.

Reference

  • GET /api/examples β€” Answer type examples (XLSForm + DDI pairs).
  • GET /api/examples/{type} β€” Single example by type (single_choice, multiple_choice, grid, integer, text, etc.).
  • GET /api/question-types β€” The registry's question-type catalogue of the pinned formtransform, keyed by qwacback's answer types: registryType, label per language (en, de), kind, base, aliases and presentation (choice: one/multiple, withOther, longList, grid, appearance). withOther, longList and appearance are the registry's. ?registry=1 returns every registry type keyed by registry slug instead. Clients get the catalogue here instead of importing @correlaid/formtransform.
  • GET /api/docs/markup-guide β€” DDI encoding conventions.

PocketBase built-in API

The studies collection is publicly readable via PocketBase's standard REST API (GET /api/collections/studies/records). Individual variables and variable_groups records are also publicly readable by ID (GET /api/collections/{name}/records/{id}), but bulk listing those collections requires authentication. Write access to all collections is admin-only.

Prefer the custom /api/questions/* endpoints over direct collection access β€” they return assembled, frontend-ready data.


Getting Started

Docker Compose (recommended)

docker compose up -d --build

Access the PocketBase Dashboard at http://localhost:8090/_/.

Data lives in seed_data/

The database is not persisted: there is no volume for /app/pb_data, locally or in production (Coolify). Every container start, and so every deploy, begins with an empty PocketBase. The migrations create the superuser from PB_ADMIN_EMAIL/PB_ADMIN_PASSWORD and import every seed_data/*.xml.

So seed_data/ is the source of truth for the question bank:

  • To add or change a study, a question or its search tags, edit the DDI file in seed_data/ and merge. The next deploy re-seeds from it.
  • Studies imported through /api/import and edits in the PocketBase dashboard last only until the next deploy.
  • Record IDs are derived from the study title and the variable/group name, so they stay the same across deploys as long as those don't change.

Published images

.github/workflows/release.yml builds qwacback and its sidecar and publishes them to GitHub Container Registry:

  • ghcr.io/correlaid/qwacback β€” the Go/PocketBase API
  • ghcr.io/correlaid/qwacback-ddi-emitter β€” the DDI ↔ XLSForm sidecar. qwacback needs it for /api/convert/*, every /xlsform endpoint, /api/examples and /api/question-types; point DDI_EMITTER_URL at it. Deploy it with the same tag as qwacback.

The validation worker (ghcr.io/correlaid/schematron-worker) is published by CorrelAid/formtransform and pulled into this stack via docker-compose.yml. The same is true for the @correlaid/formtransform library used by ddi-emitter. Both are version-pinned together via .registry-version; scripts/check-registry-version.sh (run by the release workflow) fails if they disagree.

Every push to main updates the latest tag (and a main tag). Pushing a v* git tag additionally publishes semver-pinned tags:

Git action Image tags produced
Push to main latest, main
Push tag v1.2.3 1.2.3, 1.2, 1

To cut a versioned release:

git tag v1.2.3
git push origin v1.2.3

The tag must start with v and point to a commit that includes .github/workflows/release.yml. Pin to 1.2.3 in production; use latest only for dev/staging.

Default credentials (see docker-compose.yml):

  • Admin: admin@example.com / yourpassword123
  • User: user@example.com / userpassword123

Local Development

Without validation (PocketBase only β€” imports work, validation skipped):

go run main.go serve

With validation (Docker required, since the worker is now a published image):

# Start qwacback with embedded NATS (NATS_TOKEN is required when NATS_PORT is set)
NATS_PORT=4222 NATS_TOKEN=localdev go run main.go serve &

# Pull and start the published validation worker
docker run -d --rm --name schematron-worker \
  --network host \
  -e NATS_URL=nats://localhost:4222 \
  -e NATS_TOKEN=localdev \
  ghcr.io/correlaid/schematron-worker:$(cat .registry-version)

Project Structure

internal/
  converter/    DDI ↔ XLSForm conversion, both ways via ddi-emitter
  examples/     Static answer type examples (XLSForm + DDI)
  ddixml/       Copies DDI elements as XML tokens
  exporter/     Study and question DDI, from the codebook as imported
  importer/     XML parsing β†’ PocketBase records
  routes/       API endpoints, question assembly, search
  schematron/   Go NATS client for validation worker
migrations/     Schema setup, settings, user init, seed data
ddi-emitter/    Node sidecar: DDI ↔ XLSForm via @correlaid/formtransform
seed_data/      The question bank: DDI-XML files imported on every start (the database is not persisted)
.registry-version   Pinned formtransform release tag (drives ddi-emitter and the worker image)

Development & Testing

Go Tests

Tests run during Docker build (go test in Dockerfile) and locally. No NATS or Java needed.

go test ./internal/...

Integration Test

docker compose up -d --build

# Validate only β€” no auth required
curl -X POST http://localhost:8090/api/validate \
  -F "file=@seed_data/prove_it.xml"

# Import β€” requires superuser auth
TOKEN=$(curl -s -X POST http://localhost:8090/api/collections/superusers/auth-with-password \
  -H 'Content-Type: application/json' \
  -d '{"identity":"admin@example.com","password":"yourpassword123"}' | jq -r '.token')

curl -X POST http://localhost:8090/api/import \
  -H "Authorization: Bearer $TOKEN" \
  -F "file=@seed_data/prove_it.xml"

Environment Variables

Variable Description Default
PB_ADMIN_EMAIL Initial superuser email admin@example.com
PB_ADMIN_PASSWORD Initial superuser password yourpassword123
PB_USER_EMAIL Initial regular user email user@example.com
PB_USER_PASSWORD Initial regular user password userpassword123
PB_ENCRYPTION_KEY 32-char key for settings encryption (optional)
GOMEMLIMIT Soft memory limit for Go GC 512MiB
NATS_PORT Port for embedded NATS server (optional β€” no validation without it)
NATS_TOKEN Auth token for embedded NATS server (required when NATS_PORT is set)
QWACBACK_SKIP_SEED Set to 1 to skip seeding seed_data/*.xml (the database then starts empty) (unset)

MCP Server

qwacback exposes a Model Context Protocol server at /mcp using Streamable HTTP transport. This lets AI assistants (e.g. Claude) search and browse the question bank directly.

Available tools

Tool Description
search_questions Search questions by text, concept, name, or answer type
search_studies Search studies by title, keywords, or abstract (optional topic filter)
get_question Get a single question by ID
list_questions List all questions, optionally filtered by study

All tools are read-only.

Authentication

GET /mcp (tool discovery) is public. POST /mcp and DELETE /mcp (tool calls and session teardown) require a superuser token in the Authorization: Bearer <token> header.

Client configuration

Add to your MCP client config (e.g. Claude Desktop, Claude Code):

{
  "mcpServers": {
    "qwacback": {
      "type": "streamable-http",
      "url": "http://localhost:8090/mcp",
      "headers": {
        "Authorization": "Bearer <superuser-token>"
      }
    }
  }
}

Obtain a token via:

curl -s -X POST http://localhost:8090/api/collections/superusers/auth-with-password \
  -H 'Content-Type: application/json' \
  -d '{"identity":"admin@example.com","password":"yourpassword123"}' | jq -r '.token'

Resources

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages