Rebase skyward onto upstream: tool search stops dying on stale catalogs, with rebuild and Jev logs - #7
Merged
Conversation
fforres
had a problem deploying
to
production
September 26, 2026 07:59 — with
GitHub Actions
Failure
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this is
skywardrebuilt as upstreammain(a0b0d91) plus our patch stack: 16 commits, no merge commits. It lands by force-pushingskyward-rebasedontoskyward, not with the merge button: the histories diverge on purpose, and GitHub marks this PR merged onceskywardcontains these commits. It supersedes #6, which merged upstream instead of rebasing (identical code, see below).Our changes stay visible and reviewable as a stack:
git log upstream/main..skyward.Commits (oldest first)
1-12. Our existing fork changes, replayed unchanged: delegation, D1 id pin, Access admin membership, github_app grant, email identity, User-Agent, Jev search ranking, MCP error logging (2), redeploy docs, and the two OAuth refresh lease fixes.
git range-diffshows 11 as identical patches. The OAuth lease commit differs only where its crash-loop breaker moved into upstream's rewritten stale-sync loop (UsefulSoftwareCo#2061).13.
fix(sdk): type the github_app key bytes (sdk typecheck already failed without it).14.
fix(sdk): bound background catalog rebuilds per executor. One at a time on host-cloudflare. A dead attempt retries last. The breaker joins a rebuild still running here instead of reporting it dead.15.
chore(observability): log catalog rebuild start, finish and failure, and every Jev search step (tools.list time, shard counts and sizes, each Jev call's status and duration, shard failures that were silently swallowed before).16.
chore(observability): log each step inside a rebuild (resolved, persisting with row counts and largest schema sizes, persisted).Verification
4bfdfe28…). Upstream is fully preserved: only the 46 files our commits touch differ fromupstream/main.assert-shell-assetpass.e702fb17, previousa7afa513is the rollback target). Observed in prod logs:Still open (not in this PR)
A request from one long-lived MCP client session can be cut with "session was reset" when that client's GET stream cycles while the request is in flight. The front-worker bridge WebSocket closes at the same moment. It is transport-level (the patched
agentsstreaming handler), not catalog work. It hit the Sheets and Forms explicit refreshes. The new step logs show those rebuilds were mid-write, with small rows (262 definitions, largest about 4.7k characters), when the bridge closed.