Repository navigation
regression: bench_gc_pressure ~10 -> ~17 ms on main (node 13) — was a win, now a 1.3x loss #9391
Description
Activity
Ruling out one hypothesis so nobody spends a bisect step on it:
#9359(3d1adbc4c8, "give the copied minor the liveness probes it never had") is NOT the cause.Its description reads like added per-minor work — "gives the copied minor probes it never had" — on a benchmark that collects constantly, which is a plausible mechanism and exactly why it was suggested. But the call site is gated:
if crate::gc::gc_verify_mark_enabled() { super::verify::verify_marked_heap_report_nonfatal("copying-minor"); super::verify::verify_minor_unmarked_young_children_report("copying-minor"); super::verify::verify_array_pointer_slots_enumerated_report("copying-minor"); }
PERRY_GC_VERIFY_MARKis unset in a normal benchmark run, so those probes never execute. (The probe functions are genuinely unguarded internally — they allocate aHashMapand walk every old parent on entry — so the gate is doing all the work. That is fine as-is, but it does mean a future caller that forgets the gate would be expensive; worth knowing if anyone wires them in elsewhere.)I checked this rather than accepted it because the commit came from my own fork — the direction of bias runs the wrong way for me to have taken it on trust either way.
Bisecting the remaining window now (
6bc775d9b9→9031e490dd, minus this commit). Will post the culprit.Found it, and it's mine: the concat memo from #9373. Not a bisect — the fixture named the suspect.
bench_gc_pressure's inner line is{ x: i, y: i * 2, name: "item_" + i }over 500,000 iterations. That is exactly the pattern #9373's memo targets, with every single result distinct — so the memo misses 100% of the time, inserts and evicts on every iteration, and holds up to 512 strings alive as strong GC roots on the one benchmark whose entire subject is allocation pressure. It keeps garbage alive until eviction, promotes strings that should have died young, and adds a 512-entry root scan per collection, all for zero hits.In-binary A/B (one build, a cached env kill-switch, so no build-to-build variance), min-of-15 each:
gc_pressure memo ON (as shipped) 18 ms memo OFF 12 ms 6 ms, ~50%. With the memo off this row beats node again, which matches @ECS1's "clearly winning before this window".
On my earlier suggestion of #9359 — thank you for testing it rather than taking it. It was a plausible-sounding mechanism with no evidence behind it, the gate you found rules it out, and the real cause was in my own commit in the same window. I pointed at someone else's code and the call was coming from mine.
Why my pre-merge adversarial test missed it: I did test an all-distinct-results fixture (100% miss, pure overhead) and measured 7 ms vs node's 8. But that fixture had no other allocation pressure, so it exercised the probe cost and never the part that actually hurts — rooting garbage while a collector is under load. The adversarial case has to be adversarial in the dimension the change touches, and mine only tested one of two.
Fix in progress: an admission policy, so a key must be observed twice before it earns a rooted entry. A never-repeated key then never gets inserted, which is exactly
gc_pressure, while a reused key still gets memoized on its second occurrence, which is exactlybench_object_property. Adaptive by construction — no governor, no windows, no state that can get stuck off. PR shortly; I'll post both rows' numbers.Bisected:
d923b8dcf0(#9373, the short-concat memo)Four steps over the 27-commit window, min-of-6 per build, clean
origin/mainworktree with no local changes, same host and harness throughout:commit gc_pressure e284cabc34midpoint 11 good 3fc54c4414#9367 inline transition-IC 11 good d923b8dcf0#9373 short-concat memo 17 BAD fd3c078655#9386 17 bad 9031e490ddcurrent tip 16–17 bad d923b8dcf0and its parent3fc54c4414are adjacent, so the attribution is a single commit. Node is 13–14 throughout.Note this also clears #9367 — the inline transition-IC measures 11, identical to its parent. Only the memo moves it.
Mechanism, offered as a hypothesis rather than a conclusion
The memo is a bounded 512-entry table whose entries are strong GC roots (deliberately, and for a good reason — the PR chose strong roots over pinning so a long program's distinct concats could not accumulate forever). On this benchmark that has two compounding effects, and the second is the one that matters:
- every collection traces 512 extra roots; and
- more importantly, those roots keep 512 strings alive that would otherwise be garbage, which raises the live set — and on a workload whose whole point is collecting constantly, live-set size is the cost driver, not root count.
That is the right order of magnitude for 11 → 17, but I have not confirmed it by intervention. The clean test would be a build with the memo's capacity set to 0 or its roots dropped: if
gc_pressurereturns to ~11, the mechanism is confirmed; if not, the commit is right and the mechanism is something else in it.What this is not
This is not an argument to revert #9373. It is worth 36 → 25 ms on
bench_object_propertyon its own and composes to 13 there, and that row was the board's largest loser. The question is whether the memo can keep its win without holding dead strings live across collections — a smaller capacity, eviction on collection, or weak entries with a cheap revalidation are all plausible and are the author's call, not mine.Filed by the session that hit the regression while rebasing; the memo's author has been notified directly.
- added a commit that references this issue
on Sep 1, 2026 Verified on the merged tip — the governor recovers it. Closing.
Tree
b92d7c5ac2(a tree anyone can check out), the merged label-matching harness from #9428, quiet host (load 1.8), min-of-15 per side:gc_pressure node pre-#9373 11 13 with the memo (#9373) 17 13 tip, with #9397's governor 13 12 The windowed-backoff governor (#9397) recovers the regression to within 1 ms of node — the residual ~1–2 ms over the pre-memo 11 is the governor's own accounting on a workload it correctly backs off from, which is the price of keeping the memo's
bench_object_propertywin recoverable across workload phases. Fair trade; closing.For the record, the full authoritative board from the same tree/harness/run is below — the first table of this campaign measured on a tree that actually exists, rather than on somebody's branch:
03_array_write 1 / 7 13_factorial 95 / 578 04_array_read 9 / 12 14_closure 48 / 297 05_fibonacci 339 / 972 15_mandelbrot 23 / 23 07_object_create 5 / 8 16_matrix_multiply 17 / 33 08_string_concat 5 / 27 17_loop_data_dep 220 / 219 09_method_calls 9 / 10 bench_array_grow 7 / 11 11_prime_sieve 5 / 6 bench_buffer_rw 33 / 80 12_binary_trees 6 / 9 bench_int_arith 53 / 62 02_loop_overhead 126 / 126 bench_json_ro 121 / 181 bench_json_ro_idx 134 / 179 bench_json_rt 212 / 239 bench_json_typed 212 / 238 bench_num_downgrade 3 / 4 bench_num_numeric 4 / 4 bench_string_heavy 42 / 43 bench_ta_untyped 258 / 289 bench_gc_pressure 13 / 12 * bench_object_prop 15 / 12 * 10_nested_loops 17 / 15 * 06_math_intensive 50 / 48 *25 of 29 rows at or below node. The four starred rows are 1–3 ms above, re-measured at min-of-15 to rule out the too-few-reps artifact.
bench_object_property(15 vs 12) is the only one with a candidate mechanism: it measured 13 on the memo author's branch with the always-probing memo, and the governor adds per-attempt accounting on this bench's hot path — plausible but untested, and the governor has no runtime kill switch to A/B without a build. Handing that observation to the memo's author rather than asserting it.Post-closure addendum, so the trail is complete for whoever reads the board later.
The governor that fixed this row costs 2 ms on
bench_object_property(12 → 14, node 12). That residual was bisected to the governor commit itself (5a6a64bf22, adjacent to its parent, interleaved-sealed), then chased through seven correctly-executed single-cause experiments — accounting-scheme A/B, cold-splitting the window close, a store-free gate, lazy admission tags, RMW replacement, removing the governor read entirely, and outlining the whole machinery — every one flat or worse. The outlining attempt also costgc_pressurea millisecond of call overhead on the miss path, failing the pre-stated acceptance triangle, and was reverted.The mechanism is structural, established by asm diff rather than timing: the commit grew
js_string_concat_valuefrom 759 to 851 instructions, with the growth woven through the body from the memo-attempt site onward — ~30% more global-address materialisation (adrp/movk), spill-shaped register traffic, one added atomic, one added TLS read. The effect is the sum of ten right-sized nothings, which is why no per-operation experiment could find a term: no term exists.The standing trade, measured on one host in one session each:
object_property gc_pressure no memo 21 — memo, un-governed 12 17 memo + governor (main) 14 12 Main carries one 1.17× row instead of one 1.42× row; that is the best of the three. If anyone revisits it, the honest framing is a memo redesign that never touches this function (a per-site cache), not another attempt to shave the current one.
Artifacts:
asm_0f70a367d7.txt/asm_5a6a64bf22.txt(disassembly pair), staged comparison binaries on the benchmark host. Full experiment log in the session records of the two sessions involved.Board completion: the last two starred rows are dispositioned. Every row now has a documented state.
10_nested_loops— parity by minimum: perry 16, node 16, interleaved min-of-15 in one session. The board's earlier 17/15 was session drift on both sides. The emitted inner loop does re-execute the outer-invariantarr[i]read (shadow-slot head reload + bounds check + load, where node hoists it), but at this loop's fadd-latency wall a re-load from a 24 KB L1-resident array is hidden behind the serial accumulator chain — which is also why node gains nothing from its hoist. Perry shows ~1 ms more run-to-run wobble on averages (17 vs 16); noted, not actionable.06_math_intensive— parity within clock resolution: perry 49, node 48, interleaved. The inner loop is minimal (fdiv 1.0/i; fadd; store, no poll, no helper calls) and both engines are pinned at the same fadd-latency wall, where reassociation is illegal because results must match bit-for-bit. 1 ms over 50 M iterations is 0.02 ns/iter — sub-cycle.Final board state, one line per residual:
bench_gc_pressure— recovered by fix(runtime): stop the concat memo probing when it isn't paying (from #9396) #9397's governor (12–13 vs 12–13); closed.bench_object_property— 14 vs 12; mechanism measured (asm diff: +92 instructions woven throughjs_string_concat_value), seven single-cause fixes falsified, standing trade documented above. Open path: a per-site concat cache that never enters the runtime function.10_nested_loops,06_math_intensive— parity by interleaved minimums; closed.
Every number in this thread is from a named SHA, interleaved where deltas were small, with the instruments' own failures documented alongside the results.
bench_gc_pressureregressed on main and is now above node. It was a win before.Measurements (quiet host, repeated runs, min taken)
6bc775d9b9(+ my 7 commits)origin/main@9031e490dd(clean worktree, no local changes)Stable across five runs each (first run is cold: 46/24 then settling to 16–18). Same host, same harness, same fixture.
It is not my change, and that was tested rather than argued
My branch adds a GC-poll skip for the packed-f64 loop clone (#9379), which is exactly the kind of change that could starve a collector on a GC benchmark. Ablation: with the poll-skip disabled, gc_pressure still measures 17. And plain
origin/mainwith none of my commits measures 16–17. So the cause is in main.Bisect window
That is 27 commits. I have not bisected it — GC internals are not my lane, and whoever owns the recent GC/scanner work will recognise the culprit faster than a blind bisect finds it.
Why this matters beyond one row
A regression here is not just a slow row:
gc_pressureis the workload that would catch a bad GC change, so a silent 1.3x here also degrades its value as a guard. And it poisons unrelated measurement — anyone benchmarking on main right now will attribute part of this to their own work, which is the same failure that cost a session most of a day earlier today (#9341/#9345).Caveat on my own numbers, stated because it is exactly the trap I fell into an hour ago: the "10" row was measured on a branch whose base is now 27 commits old. Same harness and host, but not the same tree — so treat the 10 as "was clearly winning before this window" rather than as a precise before-value. The 16–17 on a clean origin/main worktree with no local changes is the number I would defend.