Skip to content

Rebase skyward onto upstream: tool search stops dying on stale catalogs, with rebuild and Jev logs - #7

Merged
fforres merged 0 commit into
skywardfrom
skyward-rebased
Sep 26, 2026
Merged

fforres merged 0 commit into
skywardfrom
skyward-rebased

Conversation

@fforres

@fforres fforres commented Sep 26, 2026

Copy link
Copy Markdown

What this is

skyward rebuilt as upstream main (a0b0d91) plus our patch stack: 16 commits, no merge commits. It lands by force-pushing skyward-rebased onto skyward, not with the merge button: the histories diverge on purpose, and GitHub marks this PR merged once skyward contains these commits. It supersedes #6, which merged upstream instead of rebasing (identical code, see below).

Our changes stay visible and reviewable as a stack: git log upstream/main..skyward.

Commits (oldest first)

1-12. Our existing fork changes, replayed unchanged: delegation, D1 id pin, Access admin membership, github_app grant, email identity, User-Agent, Jev search ranking, MCP error logging (2), redeploy docs, and the two OAuth refresh lease fixes. git range-diff shows 11 as identical patches. The OAuth lease commit differs only where its crash-loop breaker moved into upstream's rewritten stale-sync loop (UsefulSoftwareCo#2061).
13. fix(sdk): type the github_app key bytes (sdk typecheck already failed without it).
14. fix(sdk): bound background catalog rebuilds per executor. One at a time on host-cloudflare. A dead attempt retries last. The breaker joins a rebuild still running here instead of reporting it dead.
15. chore(observability): log catalog rebuild start, finish and failure, and every Jev search step (tools.list time, shard counts and sizes, each Jev call's status and duration, shard failures that were silently swallowed before).
16. chore(observability): log each step inside a rebuild (resolved, persisting with row counts and largest schema sizes, persisted).

Verification

  • Same code as the merge approach: the tree at commit 14 equals Tool search stops crashing when many large catalogs are stale #6's tree exactly (4bfdfe28…). Upstream is fully preserved: only the 46 files our commits touch differ from upstream/main.
  • Tests on this branch: core/sdk 959, core/api 132, plugins/openapi 337, plugins/mcp 317 (29 skipped), plugins/toolkits 6, hosts/cloudflare 129, apps/host-cloudflare 60, execution 98. All pass. host-cloudflare build and assert-shell-asset pass.
  • Deployed from this branch (current version e702fb17, previous a7afa513 is the rollback target). Observed in prod logs:
    • Search works again, including a broad Spanish query that returns the Drive tools.
    • Background rebuilds finish one at a time (20 connections).
    • Jev is healthy: 671 candidates in 3 shards of about 57k characters, each call about 300 ms, every question answered.
    • Google Docs rebuilt on an explicit refresh.

Still open (not in this PR)

A request from one long-lived MCP client session can be cut with "session was reset" when that client's GET stream cycles while the request is in flight. The front-worker bridge WebSocket closes at the same moment. It is transport-level (the patched agents streaming handler), not catalog work. It hit the Sheets and Forms explicit refreshes. The new step logs show those rebuilds were mid-write, with small rows (262 definitions, largest about 4.7k characters), when the bridge closed.

@fforres
fforres merged commit d5126d4 into skyward Sep 26, 2026
2 checks passed

This branch had an error being deployed

1 failed deployment
production — d5126d40 Deployed Sep 26, 2026 by fforres via check #1
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant