Quick start · Features · Python clients · Documentation · Examples · Contributing
A durable Work Queue for applications that run across multiple replicas.
Workhold gives application teams reliable background execution without building queue correctness from scratch or adopting a shared platform-wide message bus. It combines idempotent intake, fenced leases, retries, follow-up tasks, and production operations behind a stable API.
- Keep work through crashes: a successful enqueue is committed to workhold's PostgreSQL before it is acknowledged.
- Scale workers safely: replicas compete for tasks under fenced, renewable leases; stale workers cannot mutate queue state.
- Finish work atomically: completion and
spawn[]follow-up tasks commit together, so the next step is not lost.
One versioned image is deployed beside each application with independent
api, migrate, maintain, relay, and apply roles. An application
business database is optional.
Run the service, migrations, and PostgreSQL locally with Docker:
git clone https://github.com/KoninMikhail/workhold-durable-work-queue.git
cd workhold-durable-work-queue
docker compose -f docker-compose.dev.yml up --buildThe public API is available on http://localhost:8080. For host development,
Python 3.13 and uv are required:
uv sync --group dev
uv run workhold --help
uv run pytestSee the local setup guide for configuration, migrations, Docker profiles, and client development.
The service is distributed as a container image. Release tags and immutable commit-SHA tags are published to GHCR. Application processes install only the role-specific client they need:
pip install workhold-producer
pip install workhold-consumer
pip install workhold-admin| Artifact | Source |
|---|---|
| Runtime image | GitHub Container Registry |
| Producer client | PyPI |
| Consumer client | PyPI |
| Admin client | PyPI |
| Wire contract | OpenAPI 3.1 |
| Release notes | GitHub Releases |
Python 3.13 or newer is required by the clients. For production deployment, read the deployment guide and security guide.
| I want to… | Start here |
|---|---|
| Understand the product in 15 minutes | Developer reading path |
| Integrate an application | Integration guide |
| See concrete workflows | Examples |
| Understand guarantees and failure behavior | Guarantees |
| Use the Python clients | Client SDK guide |
| Deploy and operate in production | Operations |
| Explore architecture and decisions | Architecture and ADRs |
| Find a specific document | Documentation index |
| Contribute a change | Contributing |
Most applications eventually need more than "put a message in a list":
- multiple replicas must not successfully own the same lease;
- retries must be bounded, delayed, and explainable;
- a timeout must not turn an enqueue retry into duplicate work;
- completing one task and creating the next must not leave partial state;
- operators need pause, drain, inspection, dead-letter, and recovery tools;
- an application should not need a business database only to get a queue.
Workhold packages those concerns into a per-application reliability boundary. It owns its PostgreSQL schema and operational state while the application continues to own payload meaning and business results.
- Named queues inside one service instance
- Idempotent enqueue with request fingerprint validation
- Priority ordering and one-shot future scheduling
- Admission control before unsafe work reaches PostgreSQL
- Atomic claim with a rotating secret token and lease generation
- Heartbeats for long-running handlers
- At-least-once recovery when a worker crashes or loses its lease
- Bounded long polling for efficient consumers
- Cooperative cancellation for already leased work
- Versioned retry policies per named queue
- Structured failure codes, delayed retries, and dead-letter state
- Idempotent terminal commands
- Attempt history and task lineage for diagnosis
complete + spawn[]in one database transaction- Follow-up tasks can target other named queues
- No state where the source succeeds but accepted follow-up work disappears
- Delivery Outbox and relay infrastructure for outbound events; event creation remains capability-gated in the public completion API
- Explicit queue creation and immutable policy versions
- Pause and drain modes
- Task, attempt, queue-depth, and dead-letter inspection
- Separate producer, worker, observer, admin, and break-glass access planes
- Migration, maintenance, retention, readiness, metrics, and audited recovery
| Use case | How workhold helps |
|---|---|
| CPU- or I/O-heavy background jobs | Distribute work across replicas, renew long leases, and safely retry after crashes |
| Multi-stage processing | Complete one task and atomically spawn the next stage into another named queue |
| Scheduled work | Make a task claimable at a bounded future available_at time without a promotion job |
| Reliable handoff from a business transaction | Use an app-local transactional outbox and deterministic idempotency key to bridge into workhold |
| Applications without a database | Enqueue directly; workhold already owns the durable store |
| Controlled operational recovery | Inspect attempts, diagnose failure codes, replay dead letters, and pause or drain queues |
| Outbound notifications | Record delivery intent in the Delivery Outbox and publish through the relay when the capability is enabled |
Explore the complete scenarios in Product use cases and Examples.
flowchart LR
producer["Producer"] -->|"idempotent enqueue"| queue["Named Work Queue"]
queue -->|"fenced claim"| worker["Worker replica"]
worker -->|"heartbeat"| queue
worker -->|"complete / fail / cancel"| queue
queue -->|"atomic spawn[]"| queue
queue -.->|"capability-gated events[]"| outbox["Delivery Outbox"]
outbox --> relay["Delivery Relay"]
relay --> destination["HTTP destination"]
Workers process tasks at least once. If a lease is lost, another worker may receive the task, so handlers and external effects must be idempotent. The service does not claim exactly-once execution or distributed transactions with an application's database. Read the guarantee matrix before integrating.
Install only the role used by each process:
pip install workhold-producer
pip install workhold-consumer
pip install workhold-admin| Package | Purpose | Optional extras |
|---|---|---|
workhold-producer |
Enqueue, inspect, cancel, and bridge | async, bridge-postgres |
workhold-consumer |
Claim, heartbeat, complete, fail, and supervise | async |
workhold-admin |
Observe, administer, and run break-glass operations | async |
The packages share a coordinated version and depend on
workhold-client-core for transport, errors, retries, and the public test
kit. OpenAPI 3.1 remains the authoritative wire contract.
- Kafka Delivery Relay adapter — publish committed Delivery Outbox events to configured Kafka topics with stable event IDs, retry handling, observability, and CloudEvents-compatible envelopes.
Kafka will be an outbound transport after commit, not a second task queue or source of truth for claims. The Work Queue lifecycle and fenced leases will remain in workhold and PostgreSQL. See ADR 018.
Workhold is a Work Queue, not a universal messaging system.
| If you need… | Use… |
|---|---|
| Pub/sub, fan-out, or event-log replay | A message broker or event log |
| DAGs, joins, compensation, or human tasks | A workflow orchestrator |
| Exactly-once external effects | Application-level idempotency, inboxes, or natural uniqueness |
| A shared multi-tenant platform bus | A platform messaging service |
RabbitMQ or Kafka can complement workhold as a downstream delivery channel; they do not need to become a second claim/ack core. See Why Workhold for the detailed comparison.
.
├── src/workhold/ # service runtime
├── packages/ # role-split Python clients
├── alembic/ # PostgreSQL migrations
├── openapi/ # OpenAPI 3.1 contract
├── tests/ # unit, integration, conformance, and chaos tests
├── benchmarks/ # qualification workloads
├── docs/ # concepts, guides, architecture, and operations
├── Dockerfile
├── docker-compose.dev.yml
└── pyproject.toml
See CONTRIBUTING.md. Workhold is a personal project.
Changes land as pull requests to main. Report a vulnerability through
SECURITY.md. AI-assisted contributors start with
AGENTS.md.
Use GitHub Issues for reproducible bugs and feature proposals. For usage questions, see SUPPORT.md. Report vulnerabilities privately according to SECURITY.md; do not open a public issue for security reports.
Workhold is available under the MIT License.