fix(ci): queue production writes FIFO so no pending run is cancelled - #147
Merged
Merged
Conversation
deploy.yml and publish-watchlist.yml share the deploy-aml-filter-com concurrency group. GitHub keeps only one pending run per group, so a third arrival (e.g. the nightly cron) silently cancelled the pending one. Each writer now first runs a credential-free `queue` job that calls the new `release-turn` Dagger function: it lists recent runs of both writers with the read-only GITHUB_TOKEN and waits until every older (lower run id) run has completed. Only then does the delivery job enter the unchanged mutex, so the group never holds a pending run to cancel. A stuck predecessor turns the run red at the 3h deadline instead of dropping it. GitHub's native `concurrency.queue: max` is the intended end state, but Foundation's pinned actionlint 1.7.10 rejects the key and the guard strips .git, so a repo-local actionlint ignore is never loaded. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015oBArfm762nN1r4F4Fst5a
hseshadr
force-pushed
the
fix/deploy-concurrency-cancel
branch
from
September 25, 2026 16:50
b908c8f to
412cafc
Compare
dagger-for-github pastes `call` raw into a bash script, so the fleet rule dagger-args-expression (hseshadr/ci#50) forbids ${{ inputs.* }}, ${{ github.event.* }} and ${{ github.head_ref }} in every input it pastes (args, call, shell, dagger-flags, workdir, cloud-token). Both delivery steps now get RELEASE_SHA through env: and pass --release-id="$RELEASE_SHA:$GITHUB_RUN_ID"; the queue step passes --run-id="$GITHUB_RUN_ID". Checkout ref keeps the same source expression (not Dagger args). Contract tests in both suites fail while any such expression is in a Dagger script input; a mutation re-adding one to publish-watchlist.yml turns both red. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015oBArfm762nN1r4F4Fst5a
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR
deploy.yml and publish-watchlist.yml share the concurrency group
deploy-aml-filter-com. GitHub keeps only one pending run per group. When a third run arrives (the nightly cron, or a second main push), GitHub cancels the pending one. The result is a lost deploy or publish that shows up as "cancelled", not red.After this PR, each writer first waits in a new
queuejob until every older deploy/publish run has finished. Only then does it enter the mutex, so the group never holds a pending run that could be cancelled. The mutex itself is unchanged.Claim touched: every green main change and every nightly list reaches production, and production writes never overlap.
Why not GitHub's native
concurrency.queue: maxThat is the official fix (GitHub changelog 2026-05-07: FIFO, up to 100 pending runs). I tried it first. It turns our CI red:
hseshadr/ci@dd19871) runs actionlint 1.7.10, which fails withunexpected key "queue" for "concurrency" section. The newest actionlint (1.7.12) fails the same way. Upstream support is still an open PR (concurrency: Add support forqueuekey rhysd/actionlint#654)..github/actionlint.yamlignore would work, but actionlint only loads it when it finds a project root, and the guard deletes.gitfrom the snapshot. I checked this with the pinned 1.7.10 image under the guard's exact invocation.Follow-up (hseshadr/ci): let the guard honor a consumer's
.github/actionlint.yaml, or bump actionlint oncequeueis supported. After that we can switch both jobs toqueue: maxand deleterelease-turn.How it works
queuejob (in both workflows): sameifguard as its delivery job. No environment, no concurrency, no secrets. Its only credential is the read-onlygithub.token(actions: readis already granted). It callsdagger call release-turn --github-token=env://GITHUB_TOKEN --run-id=${{ github.run_id }}.release-turn(.dagger/src/aml_filter/queue.py): lists the last 100 runs of both writers over HTTPS and waits while any run with a smaller run id is notcompleted. It polls every 20s and fails loudly after 3h.needs: queue. Its concurrency block is byte-identical:group: deploy-aml-filter-com,cancel-in-progress: false.Why it's safe
Evidence
test_queue.pybeforequeue.pyexistedModuleNotFoundError: aml_filter.queueruns_aheadignores statusqueues every production upload behind a credential-free turnstileneeds: queuefrom publish-watchlist.yml.daggerpoe testqueue.py100% covered; total 98.96%poe lint/typecheck(mypy) /complexity(xenon A).git-less snapshot)fetch_runs+parse_runsagainst api.github.comNot verified: the turnstile has not run inside a real production workflow. It first runs on the next main push after merge. I did not dispatch any production workflow.
Also in this PR: event values out of Dagger call text (fleet rule
dagger-args-expression, hseshadr/ci#50)dagger-for-github pastes
callraw into bash. Both delivery steps now takeRELEASE_SHAthroughenv:and pass--release-id="$RELEASE_SHA:$GITHUB_RUN_ID". The queue step passes--run-id="$GITHUB_RUN_ID". Checkoutref:keeps the same expression, since it isn't Dagger args.test_should_keep_event_expressions_out_of_dagger_script_inputs+ 4 adversarial cases, TS "keeps event expressions out…")${{ github.event.workflow_run.head_sha }}to publish-watchlist.yml → both suites red--release-id="$RELEASE_SHA:$GITHUB_RUN_ID"→ one word--release-id=<sha>:<run id>Gates:
.dagger168 passed / 1 skipped; ruff, mypy, xenon clean; TS ingress 22/22; actionlint 1.7.10 (Foundation's pin) clean.🤖 Generated with Claude Code
https://claude.ai/code/session_015oBArfm762nN1r4F4Fst5a