Skip to content

Tool search stops crashing when many large catalogs are stale - #6

Closed
fforres wants to merge 25 commits into
skywardfrom
fix/upstream-catalog-reads
Closed

fforres wants to merge 25 commits into
skywardfrom
fix/upstream-catalog-reads

Conversation

@fforres

@fforres fforres commented Sep 26, 2026

Copy link
Copy Markdown

What users see

Before this change, tool search on this host failed with "Execution lost: the session was reset" whenever many connections had stale catalogs at once, for example several large OpenAPI specs. After the crash, the unfinished-attempt backoff skipped those connections, so their catalogs never got built: a connected integration listed zero tools. With this change, search answers from the stored catalogs, and stale catalogs rebuild one at a time in the background until they all converge.

What changed

  1. Merge upstream/main into skyward (22 commits, one merge commit, no rebase). This includes Scope toolkit tool reads and stop gating reads on TTL-expired catalogs UsefulSoftwareCo/executor#2061: reads no longer wait on TTL-expired catalogs, and connection/integration lists do not wait on syncs. Two conflict hunks in packages/core/sdk/src/executor.ts were resolved by keeping both sides:
    • stampSyncedWithHealth keeps our tools_sync_started_at: null and takes upstream's new last_health shape.
    • syncStaleConnectionTools takes upstream's urgent/deferred structure, with our crash-loop breaker re-applied inside its loop.
  2. Bounded background rebuilds (c74b3a5ce). Scope toolkit tool reads and stop gating reads on TTL-expired catalogs UsefulSoftwareCo/executor#2061 alone does not fix this crash: OpenAPI catalogs have no TTL, so a stale-marked one takes the urgent path and still rebuilds 10 at a time inside the session Durable Object.
    • New ExecutorConfig.toolsSyncConcurrency, an executor-wide semaphore around background rebuilds. The SDK default stays 10; host-cloudflare sets 1.
    • The permit is taken inside the single-flight run, so a later read joins a queued rebuild instead of queueing a copy.
    • The breaker no longer treats a rebuild still running in this executor as a dead attempt.
    • A retry of an attempt that died goes to the back of its queue, so one catalog that kills its instance cannot keep the others from building.
    • Urgent rebuilds start first, so the read's grace wait is spent on the catalogs it waits for.
  3. Type-only fix in github-app.ts (Uint8Array<ArrayBuffer>). The sdk typecheck already failed without it.

No D1 schema changes.

Testing

  • New tests in connections.test.ts: with 5 stale connections, a concurrency of 1 and two overlapping reads, both reads return from the stored rows, at most one rebuild runs at a time, and each catalog rebuilds exactly once. A connection whose earlier attempt died rebuilds after the healthy one. Both tests fail with the fix removed.
  • Package tests: core/sdk 959, core/api 132, plugins/openapi 337, plugins/mcp 317 (29 skipped), plugins/toolkits 6, hosts/cloudflare 129, apps/host-cloudflare 60. All pass.
  • bun run typecheck: 44 of 45 packages pass. @executor-js/desktop fails the same way on skyward (two vite copies).
  • Lint is clean. format:check fails on the same 6 files as on skyward, none of them touched here.
  • host-cloudflare bun run build passes, including assert-shell-asset.

Risks

  • Connections already stamped by the crash stay skipped until their 15-minute backoff ends, then rebuild one at a time.
  • If a single catalog cannot fit in one session's memory on its own, it still crashes every 15 minutes, but it no longer blocks the others. Fixing that would mean moving rebuilds off the session Durable Object.
  • With a limit of 1, many stale MCP catalogs converge more slowly. A read still waits at most the 2 s grace budget.

RhysSullivan and others added 25 commits September 18, 2026 10:02
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* Align integration creation UI with admin permissions

* Show disabled integration actions for members
…UsefulSoftwareCo#2056)

Scope discovery capped the request at 100 scopes. A resource that
advertises more (PostHog lists 150) got a token missing the scopes its
MCP server needs, so every new connection synced zero tools. Bound the
request by scope-string length (8 KiB) instead.

A credential-only health check then reported healthy over the
sync-stamped rejection, hiding the failure. Sync-supplied verdicts now
carry the tool_sync_failed reason and are served until a sync succeeds.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* Accept Slack bot and user OAuth grants

* Preserve standard bearer response metadata
* Prefer OAuth when adding connections

* Verify authentication method selection stays usable

* Check client availability before preferring OAuth
…eCo#2072)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…Co#2071)

* e2e: reproduce health-probe churn on the integrations list

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Probe connection health through a per-connection atom

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* e2e: bound the churn scenario to its own connections

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…#2110) (UsefulSoftwareCo#2115)

* Revert "Test cloud admin verification and ordinary access (UsefulSoftwareCo#2110)"

This reverts commit fec546e.

* Revert "Require MFA to unlock cloud administration (UsefulSoftwareCo#2109)"

This reverts commit d0ca1b6.
…2120)

* Test organization settings MFA and ordinary access

* Verify organization settings in seat billing fixture
Brings in the 22 upstream commits, including background rebuilds for
TTL-expired catalogs and policy-scoped toolkit tool reads (UsefulSoftwareCo#2061).

Conflicts in packages/core/sdk/src/executor.ts:
- stampSyncedWithHealth: keep clearing tools_sync_started_at and take the
  upstream plugin-verdict last_health shape.
- syncStaleConnectionTools: take the upstream urgent/deferred split and
  re-apply the unfinished-attempt backoff check inside it.

Claude-Session: https://claude.ai/code/session_01VMqJkxznzTQaFVHHcttxpJ
crypto.subtle.importKey needs a BufferSource over an ArrayBuffer; the
decoded PEM was typed as Uint8Array<ArrayBufferLike>, which failed the sdk
typecheck. Both code paths already allocate a fresh ArrayBuffer.

Claude-Session: https://claude.ai/code/session_01VMqJkxznzTQaFVHHcttxpJ
A tools read could start a rebuild for every stale connection at once, up
to ten in parallel, and overlapping reads stacked more. In a Durable Object
session with several large OpenAPI catalogs this ran the instance out of
memory, the session reset, and the unfinished-attempt backoff then kept
those catalogs from ever building.

- Background rebuilds take a permit from one executor-wide semaphore,
  sized by the new toolsSyncConcurrency option (default unchanged). The
  permit is taken inside the single-flight run, so a later read joins a
  queued rebuild instead of queueing a copy.
- A start stamp from a rebuild this executor is still running is joined,
  not reported as a dead attempt.
- A retry of an attempt that never finished runs after the other stale
  catalogs, so one catalog that kills its instance cannot block the rest.
- Urgent rebuilds are started before deferred ones so the read's grace
  wait is spent on the catalogs it waits for.
- host-cloudflare runs rebuilds one at a time.

Claude-Session: https://claude.ai/code/session_01VMqJkxznzTQaFVHHcttxpJ
@fforres

fforres commented Sep 26, 2026

Copy link
Copy Markdown
Author

Superseded by #7: same code (identical tree at the bounded-rebuild commit), landed as a rebase of our patch stack onto upstream main instead of a merge, plus rebuild and Jev logging.

@fforres fforres closed this Sep 26, 2026
@fforres
fforres deleted the fix/upstream-catalog-reads branch September 26, 2026 22:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants