Repository navigation
ci: two reds on main fail every PR — gap-suite shard 5 parity regression (test_gap_10430) and a stale public benchmark baseline #10707
Description
Activity
There is a third standing red with the same shape, and it has a fix up:
e2e-scoped.It fails at "Compute e2e suite scope" after ~16 seconds on every PR:
ci_e2e_scope: these crates/perry-codegen/tests/*.rs suites are in neither SOURCE_SUITE_MAP nor SUITE_EXCLUSIONS: error_subclass_field_init, typed_collection_receiver_guardTwo suites arrived unclassified in
6925754a7d(#10443/#10446), which landed in merge train 218 (v0.5.1596) — the same train that introduced thetest_gap_10430regression in section 1 above. Like the other two, it is content-independent: #10721 (a Python script), #10722 (a.tsfixture) and #10719 (a Python script) all carry it. Fix in #10723; both suites are mapped rather than excluded, since both pass andSUITE_EXCLUSIONSis for a named failing test with an issue number.Status on the rest of this issue:
- Section 3 (build(codegen): main fails
cargo check --all-targets— two ImportedClass test initializers miss constructor_has_synthetic_arguments #10655, the twoImportedClassinitializers) is done — fix(test): unbreak the perry-codegen lib-test build on main (-D warnings --all-targets) #10682 merged at 03:55Z today and build(codegen): main failscargo check --all-targets— two ImportedClass test initializers miss constructor_has_synthetic_arguments #10655 is closed. Any PR still showingwarnings/cargo-testred was based before that and clears on rebase. - Section 1 (
test_gap_10430) is being worked. The regression window is pinned by CI evidence rather than inference: gap-suite was fully green onmainat68a5454396(train 217, 2026-09-18T17:54:22Z) and red on the very next push,60922041cd(train 218, 17:56:24Z), and red again on the 00:40 re-run of that same SHA. So the culprit is one of train 218's 30 commits, of which only 10 touch code. test_gap_2514_settracesigintis not a regression and never was. It is a recordedgap_snapshot.jsonentry (status: parity_fail, issue node:util: implement remaining top-level helper APIs #2514, added 2026-07-04), it entered when the ratchet was created on 2026-07-29 and has never left, and its fixture has not been touched since 2026-06-01. It appears in the shard's "Output Mismatches" list, but the gate'sREGRESSIONSsection names onlytest_gap_10430. Across all six shard reports of the failing run, the entire non-passing set is the 6 standing snapshot entries plustest_gap_10430— so section 1 is a single regression, not two.- Section 2 (public benchmark baseline) is unowned. Regenerating it costs ~2h on the bench mini and invalidates the committed artifact until it completes, so it needs a deliberate slot rather than an opportunistic fix.
Worth stating plainly because it is the actual thesis of this issue: train 218 is now the origin of two of the four reds. That is not a coincidence about that train so much as evidence that the per-PR tier cannot see either class — an unclassified test suite and a cross-shard parity regression are both invisible to a diff-scoped gate.
- Section 3 (build(codegen): main fails
Section 1 is stale:
test_gap_10430is already fixed. The gate is still red — for a different test, in a different shard.This supersedes the "Section 1 (
test_gap_10430) is being worked / the culprit is one of train 218's 30 commits" bullet in the previous comment. That was correct when written: the regression did land in train 218. It has since self-resolved in train 219, and I did not cause that — it was already fixed before I built anything.Both halves of section 1's framing are now wrong: the test is wrong and the shard is wrong.
test_gap_10430— six failures, then four passesAll rows are ubuntu CI on
mainexcept the last:SHA landed (UTC) test_gap_1043068a5454396train 21709-18 17:54 green (all shards) 60922041cdtrain 21809-18 17:56 FAIL — sweep runs 35377260905 and 35410163573 — PR #10646 / #10669 09-18 18:00 / 09-19 06:56 FAIL — the two runs this issue was opened from 8df83f8c1209-19 03:54 FAIL — sweep runs 35419885881 and 35423189647 b0af11e1cetrain 21909-19 08:25 pass — fast 3-shard 2 (green) and full 8-shard 7 report 4715bc2fa1train 22009-19 09:58 pass — sweep 3-shard 2 green 4715bc2fa1local— pass — 1/1, 100%, node_fail: 0, macOSSix consecutive failures then four consecutive passes is a fix, not flake.
The cleanest single datum is the train-219 full-tier shard-7 report, because at
b0af11e1ceboth tests happen to land in 8-shard 7. One build, one runner, one report:test_gap_10430_stream_module_constructor pass test_gap_array_side_mask_covers_a_pointer_stored_at_a_late_index parity_failThe handoff is visible inside a single shard, so it does not rest on comparing across runs.
Likely fix:
57506478c0(#10606), "root event/listener dispatch copies across moving GC" — flagged explicitly as UNPROVEN. Its changelog namescrates/perry-runtime/src/node_stream_event_emitter.rsas the primary site, "the path everyclass X extends EventEmittersubclass and every Node stream class (Readable,Writable,Duplex,Transform) actually dispatches through", which is exactly what the10430fixture exercises. That is good reasoning about a hypothesis I have not built, and it stays a hypothesis. I have not attributed the train-218 cause at all, and would rather leave that blank than guess.The current blocker is #10727, in shard 4
Train 219 is also where
test_gap_array_side_mask_covers_a_pointer_stored_at_a_late_indexwentpass -> parity_fail. It is absent fromgap_snapshot.json, so it blocks. Two independent observations atb0af11e1ce, in two different harness modes with different shard splits (fast 3-shard 3; auto-optimize 8-shard 7), and it passed at8df83f8c12— so it is neither flake nor a sharding artifact. Filed as #10727; no issue tracked it until now.Notably, no GC, array, or slot-enumeration file changed in train 219 at all, so the cause is not a direct code path and will need an A/B rather than more reading. Details in #10727.
Recomputing the shard, since it moves
Shard membership is round-robin over the post-filter list and shifts whenever the fixture count changes, so section 1's "shard 5" will keep going stale. To redo it:
find test-files -maxdepth 1 -type f \( -name '*.ts' -o -name '*.cts' -o -name '*.mts' \) | sort- keep entries whose basename (extension stripped) contains
test_gap_; call the 0-based positionk - shard
iofMruns exactly the tests wherek % M == i - 1(run_parity_tests.sh, theSHARD_COUNTERloop)
At
4715bc2fa1(870test_gap_fixtures), locale-independent:test index PR tier (6) sweep (3) full (8) test_gap_10430_stream_module_constructor22 5 2 7 test_gap_array_side_mask_...late_index351 4 1 8 So the blocking shard has moved 5 → 4. A PR rebased onto train 219 or later that still shows
gap-suite (5)red is showing something else, and should not be attributed to this issue.test_gap_2514_settracesigintConfirmed, matching the previous comment: a standing
gap_snapshot.jsonentry, not a regression, and it shares no cause with10430— it isutil.setTraceSigInt, untouched since 2026-06-01. Settled from the snapshot and git history without needing a run. Closing the question.Suggested disposition
Section 1 can be struck and replaced with a pointer to #10727. Sections 2 (public baseline) and the
e2e-scopeditem above are unaffected by any of this.Diagnosis for the stale-public-baseline half of this issue. It is a real regression with a specific first-bad commit, not a gate that is red by construction.
A plausible theory was tested and is false. The hypothesis was that
Cargo.tomlsits inSOURCE_PATHS, so every merge train's version bump re-invalidates the artifact — which would make the gate permanently red and any regeneration pointless within a day. The source fingerprint is byte-identical across train 225's release commit and its parent, so version bumps do not touch it.The actual first-bad commit, bisected over the fingerprint: it matches the migration target exactly at
57d3de75d3(#7958, where the migration was recorded, 2026-08-12) and diverges ate3bd92bf65— "fix(bench): reject zero-time benchmark false greens", 2026-09-01. That commit editedbenchmarks/suite/*.tsandbenchmarks/polyglot/bench.*, which are measured sources. So the published numbers were genuinely taken against older benchmark code, and the gate is reporting correctly.What that means for disposition:
- It is a real ~2 h regeneration on a quiet machine, not a structural impossibility.
- It has been standing for 18 days.
- Afterwards it stays green until the next benchmark-source edit — not until the next train. That is a much better proposition than "red forever", and it makes clearing it worthwhile rather than futile.
One concrete follow-up worth doing regardless of when the regeneration happens: the error message points at the wrong files. It reports that benchmark inputs changed, which directs a reader to
HARNESS_PATHS, the declarative config. But the harness fingerprint matches its migration target exactly — it isSOURCE_PATHSthat drifted. Anyone debugging from the message alone inspects the wrong two files and concludes nothing changed. Having the message distinguish "inputs (harness config)" from "measured sources" would have saved this bisect.Note the regeneration itself is deliberately not being done on any agent's own initiative: it rewrites the published perry-vs-node comparison, which is a claim about the project rather than a cleanup, so it is the maintainer's call.
Update: one of these is fixed, and they do not block merging
e2e-scopedis fixed on main. Both suites —error_subclass_field_initandtyped_collection_receiver_guard— are now present in_CODEGEN_SUITESinscripts/ci_e2e_scope.pyonorigin/main. They arrived unclassified with6925754a7and have since been classified. PRs still showing this failure are branched from before that fix; rebasing clears it.They do not gate merging. PR #10751 was merged today with
e2e-scopedandlintboth failing. So the practical cost of these reds is not that they stop work landing — it is the one I described above: every PR shows red for reasons unrelated to itself, so "red, but red for everyone" becomes the default reading, and a genuinely broken PR looks the same. I nearly made that error in the other direction on #10669, where a missing changelog fragment really was mine.Still open:
lint— the stale public benchmark baseline. Confirmed still failing on a PR merged today. The error names its own fix (./benchmarks/run_public_baseline.sh), and committed logs underbenchmarks/json_performance/results/carry the same message dated 2026-09-10, so this has been red for at least nine days.gap-suiteshard 5 — thetest_gap_10430_stream_module_constructorparity regression.
Neither is mine to fix blind: regenerating the public baseline commits measurement artifacts I have not reviewed the provenance of, and the gap regression wants whoever owns
node:stream/web(possibly related to #10568).- added a commit that references this issue
on Sep 20, 2026 Both halves of this issue are now resolved or relocated, so closing it.
Item 1 — gap-suite shard-5 parity regression (
test_gap_10430): fixed. Compiled and rantest_gap_10430_stream_module_constructordirectly against v0.5.1617: byte-identical to node, exit 0. The co-namedtest_gap_2514_settracesigintstill diverges, but it is a recorded entry intest-parity/gap_snapshot.json, so it cannot redden the shard.Item 2 — stale public benchmark baseline: tracked as #10799. That is the
lintstep "Public benchmark evidence freshness", still failing on today's runs, and it needs Node v22.23.1 and Bun 1.3.14 rather than a quiet machine — which is the misdiagnosis that kept it open for weeks.A note for anyone who lands here from an old link: the shard number in this issue's title went stale on its own. Shard assignment is round-robin over fixture count, so it moves whenever the fixture count changes. Record the recipe, not the index — and when picking up a long-red gate, re-read which test is failing from the current run rather than inheriting the name from the issue. This gate stayed red across a fix-and-break boundary and the continuity hid the handoff.
Main's overall CI health is tracked on #10030;
security-audit's single advisory is #10791 (unblocked by the soak window on 2026-09-21).
Two reds on main are failing every PR's CI, and neither has an owner
Found while triaging CI on #10669. Both reproduce on unrelated PRs, so they are main's, not any one change's.
1.
gap-suiteshard 5 — a real parity regressionConfirmed identical on #10669 (a GC change) and #10646 (a codegen change), with the same second mismatch (
test_gap_2514_settracesigint) in both. Shard 5 reports 140 pass / 2 parity-fail, 98.5%.Because the snapshot records it as expected-to-pass, every PR touching anything in that shard's scope now fails
gap-suite (5)and thereforepr-gate, regardless of content. That makes a real regression indistinguishable from an inherited one at a glance, which is how a genuine one gets waved through.Possibly related: #10568 (
node:stream/webReadableStream.from()returning an empty object).2.
lint— the public benchmark baseline is staleIdentical on #10643, #10651 and #10669. Every PR that touches
crates/failslinton this regardless of whether it has a changelog fragment — which also means the changeset gate's own signal is buried behind a failure nobody can act on from their own branch.Why this is worth fixing rather than routing around
Both make
pr-gatered on essentially every PR. The cost is not the red itself, it is that "CI is red, but it's red for everyone" becomes the default reading — and the next genuinely broken PR looks exactly the same. I nearly made that mistake on #10669 in the other direction: I initially assumed all of its reds were inherited, when the changelog fragment was in fact mine.A third main-health item with a fix already written is #10655 (two
ImportedClasstest initializers missingconstructor_has_synthetic_arguments, breakingcargo check --all-targets).