Skip to content

fix(ci): queue production writes FIFO so no pending run is cancelled - #147

Merged
hseshadr merged 3 commits into
mainfrom
fix/deploy-concurrency-cancel
Sep 25, 2026
Merged

hseshadr merged 3 commits into
mainfrom
fix/deploy-concurrency-cancel

Conversation

@hseshadr

@hseshadr hseshadr commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Merge after #148. This branch is stacked on feat/post-deploy-live-smoke (#148) so the two don't conflict: it contains #148's commit plus the queue job. Merge #148 first; this PR's diff then shrinks to the queue change alone. Conflict resolution kept both sides: delivery jobs are timeout-minutes: 45 (#148) and needs: queue (#147).

TL;DR

deploy.yml and publish-watchlist.yml share the concurrency group deploy-aml-filter-com. GitHub keeps only one pending run per group. When a third run arrives (the nightly cron, or a second main push), GitHub cancels the pending one. The result is a lost deploy or publish that shows up as "cancelled", not red.

After this PR, each writer first waits in a new queue job until every older deploy/publish run has finished. Only then does it enter the mutex, so the group never holds a pending run that could be cancelled. The mutex itself is unchanged.

Claim touched: every green main change and every nightly list reaches production, and production writes never overlap.

Why not GitHub's native concurrency.queue: max

That is the official fix (GitHub changelog 2026-05-07: FIFO, up to 100 pending runs). I tried it first. It turns our CI red:

  • Foundation's guard (hseshadr/ci @ dd19871) runs actionlint 1.7.10, which fails with unexpected key "queue" for "concurrency" section. The newest actionlint (1.7.12) fails the same way. Upstream support is still an open PR (concurrency: Add support for queue key rhysd/actionlint#654).
  • A repo-local .github/actionlint.yaml ignore would work, but actionlint only loads it when it finds a project root, and the guard deletes .git from the snapshot. I checked this with the pinned 1.7.10 image under the guard's exact invocation.

Follow-up (hseshadr/ci): let the guard honor a consumer's .github/actionlint.yaml, or bump actionlint once queue is supported. After that we can switch both jobs to queue: max and delete release-turn.

How it works

  • queue job (in both workflows): same if guard as its delivery job. No environment, no concurrency, no secrets. Its only credential is the read-only github.token (actions: read is already granted). It calls dagger call release-turn --github-token=env://GITHUB_TOKEN --run-id=${{ github.run_id }}.
  • release-turn (.dagger/src/aml_filter/queue.py): lists the last 100 runs of both writers over HTTPS and waits while any run with a smaller run id is not completed. It polls every 20s and fails loudly after 3h.
  • Each delivery job now has needs: queue. Its concurrency block is byte-identical: group: deploy-aml-filter-com, cancel-in-progress: false.

Why it's safe

  • Writes still never overlap. The GitHub mutex is untouched. The turnstile only controls when a run enters the group.
  • FIFO without deadlock. Run ids increase over time, so this is a strict total order. Each run waits only on smaller ids, so there can be no cycle.
  • Nothing pending to cancel. Run N+1 cannot pass the turnstile while run N is alive, so the group holds at most one run.
  • Re-runs. A re-run keeps its old, smaller id. Newer runs that have not passed the turnstile will wait for it. If a newer run is already inside the group, the mutex still serializes the two. The worst case is today's behavior, never a concurrent write.
  • Failure is loud. If GitHub can't be read or the reply is malformed, the job fails closed. A stuck predecessor turns the run red at the deadline. Nothing is silently skipped, and the token never appears in error text.
  • Known limit. Only the newest 100 runs per writer are checked. An older run still alive would need more than 100 newer runs of that writer, which won't happen at a few runs per day.

Evidence

Check Result
test_queue.py before queue.py existed RED: ModuleNotFoundError: aml_filter.queue
Mutation: runs_ahead ignores status RED: 3 tests fail
Mutation: wait on newer runs too (deadlock) RED: FIFO test fails
Workflow contracts (py) before the workflow change RED: 8 fail (turnstile gate, privileged-queue rejections, job inventory)
Ingress vitest before the workflow change RED: queues every production upload behind a credential-free turnstile
Mutation: drop needs: queue from publish-watchlist.yml RED in both the py and TS contracts
.dagger poe test 145 passed, 1 skipped; queue.py 100% covered; total 98.96%
poe lint / typecheck (mypy) / complexity (xenon A) clean
Ingress vitest + biome + tsc 20/20, clean
actionlint 1.7.10 (Foundation's pinned image, .git-less snapshot) exit 0
Read-only live fetch_runs + parse_runs against api.github.com 100 runs per writer parsed; nothing ahead

Not verified: the turnstile has not run inside a real production workflow. It first runs on the next main push after merge. I did not dispatch any production workflow.

Also in this PR: event values out of Dagger call text (fleet rule dagger-args-expression, hseshadr/ci#50)

dagger-for-github pastes call raw into bash. Both delivery steps now take RELEASE_SHA through env: and pass --release-id="$RELEASE_SHA:$GITHUB_RUN_ID". The queue step passes --run-id="$GITHUB_RUN_ID". Checkout ref: keeps the same expression, since it isn't Dagger args.

Guard Red Green
Rule tests (py test_should_keep_event_expressions_out_of_dagger_script_inputs + 4 adversarial cases, TS "keeps event expressions out…") red on the pre-migration workflows; mutation re-adding ${{ github.event.workflow_run.head_sha }} to publish-watchlist.yml → both suites red green
Bash expansion simulated the action's assemble→exec path: --release-id="$RELEASE_SHA:$GITHUB_RUN_ID" → one word --release-id=<sha>:<run id>

Gates: .dagger 168 passed / 1 skipped; ruff, mypy, xenon clean; TS ingress 22/22; actionlint 1.7.10 (Foundation's pin) clean.

🤖 Generated with Claude Code

https://claude.ai/code/session_015oBArfm762nN1r4F4Fst5a

deploy.yml and publish-watchlist.yml share the deploy-aml-filter-com
concurrency group. GitHub keeps only one pending run per group, so a
third arrival (e.g. the nightly cron) silently cancelled the pending one.

Each writer now first runs a credential-free `queue` job that calls the
new `release-turn` Dagger function: it lists recent runs of both writers
with the read-only GITHUB_TOKEN and waits until every older (lower run
id) run has completed. Only then does the delivery job enter the
unchanged mutex, so the group never holds a pending run to cancel.
A stuck predecessor turns the run red at the 3h deadline instead of
dropping it.

GitHub's native `concurrency.queue: max` is the intended end state, but
Foundation's pinned actionlint 1.7.10 rejects the key and the guard
strips .git, so a repo-local actionlint ignore is never loaded.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015oBArfm762nN1r4F4Fst5a
@hseshadr
hseshadr force-pushed the fix/deploy-concurrency-cancel branch from b908c8f to 412cafc Compare September 25, 2026 16:50
dagger-for-github pastes `call` raw into a bash script, so the fleet rule
dagger-args-expression (hseshadr/ci#50) forbids ${{ inputs.* }},
${{ github.event.* }} and ${{ github.head_ref }} in every input it pastes
(args, call, shell, dagger-flags, workdir, cloud-token). Both delivery steps
now get RELEASE_SHA through env: and pass --release-id="$RELEASE_SHA:$GITHUB_RUN_ID";
the queue step passes --run-id="$GITHUB_RUN_ID". Checkout ref keeps the same
source expression (not Dagger args).

Contract tests in both suites fail while any such expression is in a Dagger
script input; a mutation re-adding one to publish-watchlist.yml turns both red.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015oBArfm762nN1r4F4Fst5a
@hseshadr
hseshadr merged commit 4392391 into main Sep 25, 2026
2 checks passed
@hseshadr
hseshadr deleted the fix/deploy-concurrency-cancel branch September 25, 2026 19:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant