Skip to content

feat(observability): OpenTelemetry traces and metrics - #49

Merged
PhilippTheServer merged 1 commit into
OpenTaberna:mainfrom
PhilippTheServer:pl/otel
Aug 26, 2026
Merged

PhilippTheServer merged 1 commit into
OpenTaberna:mainfrom
PhilippTheServer:pl/otel

Conversation

@PhilippTheServer

Copy link
Copy Markdown
Contributor

Closes #48 · S3 of the telemetry programme

Traces and metrics from the API and the worker, over OTLP, with a self-hosted collector,
Prometheus and Grafana in the dev compose file.

API ────┐
        ├──▶ OTel Collector ──▶ Prometheus ──▶ Grafana
Worker ─┘         (:4318)         (:9090)      (:3001)

OTLP is the seam

No application module imports a vendor SDK. Using Datadog or Grafana Cloud is a change to
OTEL_EXPORTER_OTLP_ENDPOINT and the collector config, not to any service.

Off by default — while OTEL_ENABLED is false, setup() returns before creating an
exporter, so a deployment that has not opted in opens no connection.

Every step is wrapped: an absent or misconfigured collector produces a warning and a running
API. Observability that can cause the outage it exists to diagnose is a bad trade.

The gauges Deployment.md told operators to watch

They were queries someone had to remember to run, which means nobody ran them and the first
sign of a stalled pipeline was a customer asking where their parcel was.

Metric Non-zero means
opentaberna.outbox.pending Rising: the worker is not running
opentaberna.outbox.failed Never reached the queue — Redis or the poller
opentaberna.outbox.dead Jobs ran and gave up — usually the carrier
opentaberna.webhooks.unprocessed Payments arriving, not handled. Page on this.
opentaberna.orders.awaiting_shipment The work queue

Two bugs, both found by testing the pipeline's output rather than its setup

instrument_app ran before telemetry was configured. main.py executes at import, long
before lifespan startup, so the guard turned instrumentation into a no-op and no HTTP
metrics were produced at all
— while startup logged OpenTelemetry configured throughout.
A test asserting "setup was called" would have passed.

Business gauges silently produced nothing. They were observable gauges, whose callbacks
run on the metrics SDK's own thread, and the only engine here is asynchronous. Every
collection failed. Moved to a 30-second cron in the worker, which already has a scheduler
and an async session.

Both are the worst kind of monitoring bug: the thing reports itself healthy and emits
nothing.

Verification

$ uv run pytest tests/test_telemetry_unit.py tests/test_telemetry_integration.py -q
20 passed

$ uv run pytest -q
1010 passed, 5 skipped in 17.91s

$ ruff check src/app tests/
All checks passed!

Against the live stack — 11/11 dashboard panels return data:

OK  Outbox pending (1)          OK  Request rate by route (2)
OK  Outbox failed (1)           OK  Error rate (5xx) (1)
OK  Outbox dead (1)             OK  Latency p95 by route (2)
OK  Webhooks unprocessed (1)    OK  Latency p99 (1)
OK  Awaiting shipment (1)       OK  Database connection pool (2)
                                OK  Requests in flight (1)

opentaberna_orders_awaiting_shipment = 13, cross-checked against the SQL it claims to
represent. Prometheus target opentaberna is up. Both job="opentaberna-api" and
job="opentaberna-worker" report.

Three dashboard details worth keeping

The error-rate panel ends in or vector(0). Without it a healthy shop shows "No data",
indistinguishable from a broken scrape — the wrong thing to be uncertain about mid-incident.

Latency is p95/p99, never an average. An average hides the tail customers notice.

Panels group by http_target, not http_route. The FastAPI instrumentation uses the
former; grouping by the latter silently collapses every route into one unnamed series, which
looks like a working panel. I made that mistake and caught it by checking each panel
returned more than one series.

Grafana is anonymous in the dev compose file only, and the docs say not to copy that.

The API had structured logs and correlation IDs and nothing else. "The shop
feels slow" could not be answered with anything but a guess, and a regression in
one endpoint stayed invisible until someone reported it. Deployment.md already
named the queue states worth alerting on, and nothing exported them, so nothing
could alert.

Adds OTel traces and metrics from the API and the worker, covering HTTP,
SQLAlchemy, Redis and outbound HTTP, exported over OTLP. The dev compose file
gains a collector, Prometheus and Grafana with a provisioned dashboard.

OTLP is the seam. No application module imports a vendor SDK, so using Datadog
or Grafana Cloud is a change to OTEL_EXPORTER_OTLP_ENDPOINT and the collector
config rather than to any service.

Off by default: while OTEL_ENABLED is false, setup() returns before creating an
exporter, so a deployment that has not opted in opens no connection.

Every step of the wiring is wrapped. An absent or misconfigured collector
produces a warning and a running API, not a failed start — observability that
can cause the outage it exists to diagnose is a bad trade.

Health endpoints are excluded from tracing. A liveness probe every few seconds
would bury real traffic under a heartbeat.

Log records now carry trace_id beside correlation_id. Without it the two systems
describe the same request and cannot be put side by side, so a slow span in
Grafana leads nowhere.

Two bugs found by testing the pipeline's output rather than its setup:

instrument_app ran at module import, long before lifespan configured telemetry,
so the guard turned it into a no-op and no HTTP metrics were produced at all.
Startup logged cleanly throughout. setup() now runs in main.py before
instrumenting, and remains idempotent for lifespan's call.

Business gauges were observable gauges whose callbacks run on the metrics SDK's
own thread, and the only engine here is asynchronous — so collection failed
every time and the gauges registered cleanly while producing nothing. Moved to a
30-second cron in the worker, which already has a scheduler and an async
session, and is the process that most needs to be alive for those numbers to
matter.

Closes #48

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YY1ekLLeFLkAU2kvdQ8Ey4
@PhilippTheServer
PhilippTheServer merged commit 49dc71a into OpenTaberna:main Aug 26, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

No metrics, no traces: a slow endpoint is invisible until a customer complains

1 participant