Skip to content

feat(telemetry): capture uncaught frontend errors - #51

Merged
PhilippTheServer merged 1 commit into
OpenTaberna:mainfrom
PhilippTheServer:pl/frontend-errors
Aug 26, 2026
Merged

PhilippTheServer merged 1 commit into
OpenTaberna:mainfrom
PhilippTheServer:pl/frontend-errors

Conversation

@PhilippTheServer

Copy link
Copy Markdown
Contributor

Closes #50 · S4, the last part of the telemetry programme

S3 gave the API traces and metrics. Neither sees a component throwing in somebody's browser:
the server returns 200, the metrics look healthy, and the shop is broken for a real customer.

POST /v1/telemetry/errors        public, rate limited, opt-in
GET  /v1/admin/telemetry/errors  grouped by app + class + message, by frequency

The user agent is reduced, never stored

A raw agent string is a fingerprint. But "which browser?" is genuinely diagnostic — a large
share of frontend bugs are one engine behaving differently — so throwing it away entirely
makes the reports much weaker.

It is reduced at the boundary to a family and major version, and only that is stored.
Safari 18 reproduces a bug; it does not recognise anyone. The reduction doubles as a
filter — whatever a client sends, the output is a known family name and an integer:

"Mozilla/5.0 Chrome/140 user=alice@example.com token=abc123"  →  "Chrome 140"

Edge and Opera both claim to be Chrome, and Chrome claims to be Safari, so the patterns are
ordered most-specific-first. Without that every error would look like it came from Chrome —
there is a test for it.

Grouped by message, not by stack

One bug produces thousands of identical rows. Grouping by stack would split a single fault
reached from two routes into two bugs, so affected_paths carries the spread and one
representative stack comes back for debugging.

Public, and treated as such

Rate limited to 30/min — tighter than the analytics ingest, because a component throwing in
a render loop reports as fast as the browser can loop. Batch capped at 10, closed app
vocabulary, extra="forbid", every field bounded, stacks truncated at 4000 characters.

What it cannot tell you

It reports only what browsers managed to send. An error that breaks a page badly enough to
stop the reporter is precisely the one that will not appear. Silence means no news, not no
errors — and the docs say so rather than letting a reader assume otherwise.

Verification

$ uv run pytest tests/test_frontend_errors_unit.py tests/test_frontend_errors_integration.py -q
22 passed

$ uv run pytest -q
1048 passed, 7 skipped in 17.63s

$ ruff check src/app tests/
All checks passed!

Live against the stack:

report as Safari, 4 errors incl. one dated 1999  →  accepted: 3  rejected: 1
stored:  storefront | TypeError | /shop/1 | Safari 18     ← "?token=secret" stripped
         storefront | TypeError | /shop/2 | Safari 18
         admin      | HttpErrorResponse | /orders | Safari 18

grouped: [storefront] TypeError x2  paths=2  browsers=['Safari 18']
         [admin] HttpErrorResponse x1  paths=1  browsers=['Safari 18']

admin read without a token → 403

The Angular reporters in both frontends land separately.

Traces and metrics see nothing that happens in a browser. A component that
throws leaves the server returning 200 with healthy metrics while the shop is
broken for a real customer — and for a shop, by the time somebody reports it the
sale is already lost.

Adds a public report endpoint and an admin read that groups errors by
application, class and message.

    POST /v1/telemetry/errors        public, rate limited, opt-in
    GET  /v1/admin/telemetry/errors  grouped, ordered by frequency

The report endpoint takes no token because storefront visitors are not signed in
and an error before login is exactly the one worth catching. That makes it
hostile-input territory, so it is rate limited harder than the analytics ingest
— a component throwing inside a render loop is the normal failure mode here and
reports as fast as the browser can loop — with a closed app vocabulary, forbidden
extra fields and every field bounded. Stacks are truncated rather than rejected:
unbounded input from a public endpoint, but also the most useful field, and the
top frames are where the fault is.

The user agent is reduced at the boundary to a family and major version and only
that is stored. A raw agent string is a fingerprint, but "which browser" is
genuinely diagnostic, so discarding it entirely would make the reports much
weaker. "Safari 18" reproduces a bug and does not recognise anyone. The
reduction doubles as a filter: whatever a client sends, the output is a known
family name and an integer, never a fragment of the input.

Errors are grouped by message rather than by stack. The same fault reached from
two routes produces two stacks and is one bug; affected_paths shows the spread
instead.

Documented plainly: this reports only what browsers managed to send, so an error
that breaks a page badly enough to stop the reporter is the one that will not
appear. Silence means no news, not no errors.

Closes #50

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YY1ekLLeFLkAU2kvdQ8Ey4
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A broken storefront is invisible: JavaScript errors reach nobody

1 participant