Skip to content

fix(events): dedupe SSE emission, scope live streams, fix reconnect leak - #26

Merged
iaj6 merged 1 commit into
mainfrom
fix/sse-cluster
Jul 12, 2026
Merged

iaj6 merged 1 commit into
mainfrom
fix/sse-cluster

Conversation

@iaj6

@iaj6 iaj6 commented Jul 12, 2026

Copy link
Copy Markdown
Owner

Summary

Fixes the SSE/event-streaming bug cluster from the July audit's medium tier:

  • Exactly-once event emission — the SSE route polled with an inclusive since timestamp, so the newest event (and any sharing its timestamp) was re-enqueued every 2s forever. New shared EventPollCursor in @agentops/db tracks seen IDs at the boundary timestamp: no re-emits, and a late arrival sharing an already-emitted timestamp still gets delivered exactly once (a bare gt would drop it).
  • Events page "N total" no longer inflates — the counter only increments when an event is actually newly inserted (pure applyLiveEvent reducer, unit-tested).
  • Reconnect-timer leak fixed — the backoff setTimeout handle is stored and cleared on unmount/filter change, so a pending reconnect can't open an orphan EventSource (each leaked stream kept a server-side 2s poll + 15s ping alive).
  • Live streams respect view scope — useRuns/useEvents now forward the admin's view/userId scope to the SSE URL (previously the filter silently stopped applying the moment anything streamed), with a client-side guard in useRuns as belt-and-suspenders.
  • Ownership set no longer frozen at connect — the route re-resolves the user's owned run/session IDs on each poll, so runs created after connecting stream to their owner immediately instead of after a reconnect.
  • agentops events tail reuses the same cursor, fixing its identical reprint-every-second bug.

Tests

Cursor pure-function + real-SQLite integration tests (incl. shared-timestamp insert between polls), an explicit test pinning listEvents' inclusive-since semantics the cursor relies on, and reducer tests for the total counter. No React-rendering tests — the repo has no jsdom setup, so hook changes are covered via extracted pure logic (noted in-code). Full workspace suite green (1,302 tests), web lint clean.

Known pre-existing gap left as-is: a poll reads at most 100 events, so a >100-event burst between polls can skip the overflow.

🤖 Generated with Claude Code

Fixes the July 2026 audit's SSE/event-streaming cluster:

- Add an EventPollCursor to @agentops/db pairing an inclusive (gte)
  since boundary with the event IDs already emitted at that timestamp,
  so polling emits each event exactly once without dropping a second
  event that shares the boundary timestamp. Used by both the web SSE
  route and CLI 'events tail', which previously re-emitted the newest
  event on every poll.
- useEvents: hold events + total in one state object updated through a
  pure applyLiveEvent reducer, so the 'N total' counter only increments
  when an event is genuinely new instead of on every SSE delivery.
- useEventSource: store the backoff reconnect timer handle and clear it
  on cleanup and on (re)connect, so unmount/filter changes can no longer
  leak a pending reconnect that opens an EventSource nobody closes.
- Forward view/userId scope params on the SSE connection from useRuns
  and useEvents (plus a client-side run.userId guard in useRuns), so an
  admin viewing ?userId=X no longer sees other users' live runs/events.
- SSE route: re-resolve the owned sourceId set on each poll so runs and
  sessions created after connect appear in the owner's live stream
  without a reconnect.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@iaj6
iaj6 merged commit d55e720 into main Jul 12, 2026
3 checks passed
@iaj6
iaj6 deleted the fix/sse-cluster branch July 12, 2026 15:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant