Repository navigation
feat(observability): OpenTelemetry traces and metrics - #49
Merged
Merged
Conversation
The API had structured logs and correlation IDs and nothing else. "The shop feels slow" could not be answered with anything but a guess, and a regression in one endpoint stayed invisible until someone reported it. Deployment.md already named the queue states worth alerting on, and nothing exported them, so nothing could alert. Adds OTel traces and metrics from the API and the worker, covering HTTP, SQLAlchemy, Redis and outbound HTTP, exported over OTLP. The dev compose file gains a collector, Prometheus and Grafana with a provisioned dashboard. OTLP is the seam. No application module imports a vendor SDK, so using Datadog or Grafana Cloud is a change to OTEL_EXPORTER_OTLP_ENDPOINT and the collector config rather than to any service. Off by default: while OTEL_ENABLED is false, setup() returns before creating an exporter, so a deployment that has not opted in opens no connection. Every step of the wiring is wrapped. An absent or misconfigured collector produces a warning and a running API, not a failed start — observability that can cause the outage it exists to diagnose is a bad trade. Health endpoints are excluded from tracing. A liveness probe every few seconds would bury real traffic under a heartbeat. Log records now carry trace_id beside correlation_id. Without it the two systems describe the same request and cannot be put side by side, so a slow span in Grafana leads nowhere. Two bugs found by testing the pipeline's output rather than its setup: instrument_app ran at module import, long before lifespan configured telemetry, so the guard turned it into a no-op and no HTTP metrics were produced at all. Startup logged cleanly throughout. setup() now runs in main.py before instrumenting, and remains idempotent for lifespan's call. Business gauges were observable gauges whose callbacks run on the metrics SDK's own thread, and the only engine here is asynchronous — so collection failed every time and the gauges registered cleanly while producing nothing. Moved to a 30-second cron in the worker, which already has a scheduler and an async session, and is the process that most needs to be alive for those numbers to matter. Closes #48 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YY1ekLLeFLkAU2kvdQ8Ey4
PhilippTheServer
force-pushed
the
pl/otel
branch
from
August 26, 2026 14:53
1972f39 to
8afb497
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #48 · S3 of the telemetry programme
Traces and metrics from the API and the worker, over OTLP, with a self-hosted collector,
Prometheus and Grafana in the dev compose file.
OTLP is the seam
No application module imports a vendor SDK. Using Datadog or Grafana Cloud is a change to
OTEL_EXPORTER_OTLP_ENDPOINTand the collector config, not to any service.Off by default — while
OTEL_ENABLEDis false,setup()returns before creating anexporter, so a deployment that has not opted in opens no connection.
Every step is wrapped: an absent or misconfigured collector produces a warning and a running
API. Observability that can cause the outage it exists to diagnose is a bad trade.
The gauges Deployment.md told operators to watch
They were queries someone had to remember to run, which means nobody ran them and the first
sign of a stalled pipeline was a customer asking where their parcel was.
opentaberna.outbox.pendingopentaberna.outbox.failedopentaberna.outbox.deadopentaberna.webhooks.unprocessedopentaberna.orders.awaiting_shipmentTwo bugs, both found by testing the pipeline's output rather than its setup
instrument_appran before telemetry was configured.main.pyexecutes at import, longbefore lifespan startup, so the guard turned instrumentation into a no-op and no HTTP
metrics were produced at all — while startup logged
OpenTelemetry configuredthroughout.A test asserting "setup was called" would have passed.
Business gauges silently produced nothing. They were observable gauges, whose callbacks
run on the metrics SDK's own thread, and the only engine here is asynchronous. Every
collection failed. Moved to a 30-second cron in the worker, which already has a scheduler
and an async session.
Both are the worst kind of monitoring bug: the thing reports itself healthy and emits
nothing.
Verification
Against the live stack — 11/11 dashboard panels return data:
opentaberna_orders_awaiting_shipment= 13, cross-checked against the SQL it claims torepresent. Prometheus target
opentabernaisup. Bothjob="opentaberna-api"andjob="opentaberna-worker"report.Three dashboard details worth keeping
The error-rate panel ends in
or vector(0). Without it a healthy shop shows "No data",indistinguishable from a broken scrape — the wrong thing to be uncertain about mid-incident.
Latency is p95/p99, never an average. An average hides the tail customers notice.
Panels group by
http_target, nothttp_route. The FastAPI instrumentation uses theformer; grouping by the latter silently collapses every route into one unnamed series, which
looks like a working panel. I made that mistake and caught it by checking each panel
returned more than one series.
Grafana is anonymous in the dev compose file only, and the docs say not to copy that.