Skip to content

regression: bench_gc_pressure ~10 -> ~17 ms on main (node 13) — was a win, now a 1.3x loss #9391

Description

@proggeramlug

bench_gc_pressure regressed on main and is now above node. It was a win before.

Measurements (quiet host, repeated runs, min taken)

tree gc_pressure node
my branch based on 6bc775d9b9 (+ my 7 commits) 10 13
plain origin/main @ 9031e490dd (clean worktree, no local changes) 16–17 13
my branch rebased onto current main 17 13

Stable across five runs each (first run is cold: 46/24 then settling to 16–18). Same host, same harness, same fixture.

It is not my change, and that was tested rather than argued

My branch adds a GC-poll skip for the packed-f64 loop clone (#9379), which is exactly the kind of change that could starve a collector on a GC benchmark. Ablation: with the poll-skip disabled, gc_pressure still measures 17. And plain origin/main with none of my commits measures 16–17. So the cause is in main.

Bisect window

6bc775d9b9 refactor(codegen): split the ws dispatch rows out of net_events.rs (#9343)
  ..
757beace0f fix: symbol-filter isolation (#9383) + CONCAT_MEMO thread-local (#9390)

That is 27 commits. I have not bisected it — GC internals are not my lane, and whoever owns the recent GC/scanner work will recognise the culprit faster than a blind bisect finds it.

Why this matters beyond one row

A regression here is not just a slow row: gc_pressure is the workload that would catch a bad GC change, so a silent 1.3x here also degrades its value as a guard. And it poisons unrelated measurement — anyone benchmarking on main right now will attribute part of this to their own work, which is the same failure that cost a session most of a day earlier today (#9341/#9345).

Caveat on my own numbers, stated because it is exactly the trap I fell into an hour ago: the "10" row was measured on a branch whose base is now 27 commits old. Same harness and host, but not the same tree — so treat the 10 as "was clearly winning before this window" rather than as a precise before-value. The 16–17 on a clean origin/main worktree with no local changes is the number I would defend.

Activity

  1. proggeramlug commented on Sep 1, 2026

    @proggeramlug
    ContributorAuthor

    Ruling out one hypothesis so nobody spends a bisect step on it: #9359 (3d1adbc4c8, "give the copied minor the liveness probes it never had") is NOT the cause.

    Its description reads like added per-minor work — "gives the copied minor probes it never had" — on a benchmark that collects constantly, which is a plausible mechanism and exactly why it was suggested. But the call site is gated:

    if crate::gc::gc_verify_mark_enabled() {
        super::verify::verify_marked_heap_report_nonfatal("copying-minor");
        super::verify::verify_minor_unmarked_young_children_report("copying-minor");
        super::verify::verify_array_pointer_slots_enumerated_report("copying-minor");
    }

    PERRY_GC_VERIFY_MARK is unset in a normal benchmark run, so those probes never execute. (The probe functions are genuinely unguarded internally — they allocate a HashMap and walk every old parent on entry — so the gate is doing all the work. That is fine as-is, but it does mean a future caller that forgets the gate would be expensive; worth knowing if anyone wires them in elsewhere.)

    I checked this rather than accepted it because the commit came from my own fork — the direction of bias runs the wrong way for me to have taken it on trust either way.

    Bisecting the remaining window now (6bc775d9b9 → 9031e490dd, minus this commit). Will post the culprit.

  2. proggeramlug commented on Sep 1, 2026

    @proggeramlug
    ContributorAuthor

    Found it, and it's mine: the concat memo from #9373. Not a bisect — the fixture named the suspect.

    bench_gc_pressure's inner line is { x: i, y: i * 2, name: "item_" + i } over 500,000 iterations. That is exactly the pattern #9373's memo targets, with every single result distinct — so the memo misses 100% of the time, inserts and evicts on every iteration, and holds up to 512 strings alive as strong GC roots on the one benchmark whose entire subject is allocation pressure. It keeps garbage alive until eviction, promotes strings that should have died young, and adds a 512-entry root scan per collection, all for zero hits.

    In-binary A/B (one build, a cached env kill-switch, so no build-to-build variance), min-of-15 each:

    gc_pressure
    memo ON (as shipped) 18 ms
    memo OFF 12 ms

    6 ms, ~50%. With the memo off this row beats node again, which matches @ECS1's "clearly winning before this window".

    On my earlier suggestion of #9359 — thank you for testing it rather than taking it. It was a plausible-sounding mechanism with no evidence behind it, the gate you found rules it out, and the real cause was in my own commit in the same window. I pointed at someone else's code and the call was coming from mine.

    Why my pre-merge adversarial test missed it: I did test an all-distinct-results fixture (100% miss, pure overhead) and measured 7 ms vs node's 8. But that fixture had no other allocation pressure, so it exercised the probe cost and never the part that actually hurts — rooting garbage while a collector is under load. The adversarial case has to be adversarial in the dimension the change touches, and mine only tested one of two.

    Fix in progress: an admission policy, so a key must be observed twice before it earns a rooted entry. A never-repeated key then never gets inserted, which is exactly gc_pressure, while a reused key still gets memoized on its second occurrence, which is exactly bench_object_property. Adaptive by construction — no governor, no windows, no state that can get stuck off. PR shortly; I'll post both rows' numbers.

  3. proggeramlug commented on Sep 1, 2026

    @proggeramlug
    ContributorAuthor

    Bisected: d923b8dcf0 (#9373, the short-concat memo)

    Four steps over the 27-commit window, min-of-6 per build, clean origin/main worktree with no local changes, same host and harness throughout:

    commit gc_pressure
    e284cabc34 midpoint 11 good
    3fc54c4414 #9367 inline transition-IC 11 good
    d923b8dcf0 #9373 short-concat memo 17 BAD
    fd3c078655 #9386 17 bad
    9031e490dd current tip 16–17 bad

    d923b8dcf0 and its parent 3fc54c4414 are adjacent, so the attribution is a single commit. Node is 13–14 throughout.

    Note this also clears #9367 — the inline transition-IC measures 11, identical to its parent. Only the memo moves it.

    Mechanism, offered as a hypothesis rather than a conclusion

    The memo is a bounded 512-entry table whose entries are strong GC roots (deliberately, and for a good reason — the PR chose strong roots over pinning so a long program's distinct concats could not accumulate forever). On this benchmark that has two compounding effects, and the second is the one that matters:

    1. every collection traces 512 extra roots; and
    2. more importantly, those roots keep 512 strings alive that would otherwise be garbage, which raises the live set — and on a workload whose whole point is collecting constantly, live-set size is the cost driver, not root count.

    That is the right order of magnitude for 11 → 17, but I have not confirmed it by intervention. The clean test would be a build with the memo's capacity set to 0 or its roots dropped: if gc_pressure returns to ~11, the mechanism is confirmed; if not, the commit is right and the mechanism is something else in it.

    What this is not

    This is not an argument to revert #9373. It is worth 36 → 25 ms on bench_object_property on its own and composes to 13 there, and that row was the board's largest loser. The question is whether the memo can keep its win without holding dead strings live across collections — a smaller capacity, eviction on collection, or weak entries with a cheap revalidation are all plausible and are the author's call, not mine.

    Filed by the session that hit the regression while rebasing; the memo's author has been notified directly.

  4. proggeramlug commented on Sep 2, 2026

    @proggeramlug
    ContributorAuthor

    Verified on the merged tip — the governor recovers it. Closing.

    Tree b92d7c5ac2 (a tree anyone can check out), the merged label-matching harness from #9428, quiet host (load 1.8), min-of-15 per side:

    gc_pressure node
    pre-#9373 11 13
    with the memo (#9373) 17 13
    tip, with #9397's governor 13 12

    The windowed-backoff governor (#9397) recovers the regression to within 1 ms of node — the residual ~1–2 ms over the pre-memo 11 is the governor's own accounting on a workload it correctly backs off from, which is the price of keeping the memo's bench_object_property win recoverable across workload phases. Fair trade; closing.

    For the record, the full authoritative board from the same tree/harness/run is below — the first table of this campaign measured on a tree that actually exists, rather than on somebody's branch:

    03_array_write        1 /   7    13_factorial          95 / 578
    04_array_read         9 /  12    14_closure            48 / 297
    05_fibonacci        339 / 972    15_mandelbrot         23 /  23
    07_object_create      5 /   8    16_matrix_multiply    17 /  33
    08_string_concat      5 /  27    17_loop_data_dep     220 / 219
    09_method_calls       9 /  10    bench_array_grow       7 /  11
    11_prime_sieve        5 /   6    bench_buffer_rw       33 /  80
    12_binary_trees       6 /   9    bench_int_arith       53 /  62
    02_loop_overhead    126 / 126    bench_json_ro        121 / 181
    bench_json_ro_idx   134 / 179    bench_json_rt        212 / 239
    bench_json_typed    212 / 238    bench_num_downgrade    3 /   4
    bench_num_numeric     4 /   4    bench_string_heavy    42 /  43
    bench_ta_untyped    258 / 289    bench_gc_pressure     13 /  12 *
    bench_object_prop    15 /  12 *  10_nested_loops       17 /  15 *
    06_math_intensive    50 /  48 *
    

    25 of 29 rows at or below node. The four starred rows are 1–3 ms above, re-measured at min-of-15 to rule out the too-few-reps artifact. bench_object_property (15 vs 12) is the only one with a candidate mechanism: it measured 13 on the memo author's branch with the always-probing memo, and the governor adds per-attempt accounting on this bench's hot path — plausible but untested, and the governor has no runtime kill switch to A/B without a build. Handing that observation to the memo's author rather than asserting it.

  5. proggeramlug commented on Sep 2, 2026

    @proggeramlug
    ContributorAuthor

    Post-closure addendum, so the trail is complete for whoever reads the board later.

    The governor that fixed this row costs 2 ms on bench_object_property (12 → 14, node 12). That residual was bisected to the governor commit itself (5a6a64bf22, adjacent to its parent, interleaved-sealed), then chased through seven correctly-executed single-cause experiments — accounting-scheme A/B, cold-splitting the window close, a store-free gate, lazy admission tags, RMW replacement, removing the governor read entirely, and outlining the whole machinery — every one flat or worse. The outlining attempt also cost gc_pressure a millisecond of call overhead on the miss path, failing the pre-stated acceptance triangle, and was reverted.

    The mechanism is structural, established by asm diff rather than timing: the commit grew js_string_concat_value from 759 to 851 instructions, with the growth woven through the body from the memo-attempt site onward — ~30% more global-address materialisation (adrp/movk), spill-shaped register traffic, one added atomic, one added TLS read. The effect is the sum of ten right-sized nothings, which is why no per-operation experiment could find a term: no term exists.

    The standing trade, measured on one host in one session each:

    object_property gc_pressure
    no memo 21 —
    memo, un-governed 12 17
    memo + governor (main) 14 12

    Main carries one 1.17× row instead of one 1.42× row; that is the best of the three. If anyone revisits it, the honest framing is a memo redesign that never touches this function (a per-site cache), not another attempt to shave the current one.

    Artifacts: asm_0f70a367d7.txt / asm_5a6a64bf22.txt (disassembly pair), staged comparison binaries on the benchmark host. Full experiment log in the session records of the two sessions involved.

  6. proggeramlug commented on Sep 2, 2026

    @proggeramlug
    ContributorAuthor

    Board completion: the last two starred rows are dispositioned. Every row now has a documented state.

    10_nested_loops — parity by minimum: perry 16, node 16, interleaved min-of-15 in one session. The board's earlier 17/15 was session drift on both sides. The emitted inner loop does re-execute the outer-invariant arr[i] read (shadow-slot head reload + bounds check + load, where node hoists it), but at this loop's fadd-latency wall a re-load from a 24 KB L1-resident array is hidden behind the serial accumulator chain — which is also why node gains nothing from its hoist. Perry shows ~1 ms more run-to-run wobble on averages (17 vs 16); noted, not actionable.

    06_math_intensive — parity within clock resolution: perry 49, node 48, interleaved. The inner loop is minimal (fdiv 1.0/i; fadd; store, no poll, no helper calls) and both engines are pinned at the same fadd-latency wall, where reassociation is illegal because results must match bit-for-bit. 1 ms over 50 M iterations is 0.02 ns/iter — sub-cycle.

    Final board state, one line per residual:

    • bench_gc_pressure — recovered by fix(runtime): stop the concat memo probing when it isn't paying (from #9396) #9397's governor (12–13 vs 12–13); closed.
    • bench_object_property — 14 vs 12; mechanism measured (asm diff: +92 instructions woven through js_string_concat_value), seven single-cause fixes falsified, standing trade documented above. Open path: a per-site concat cache that never enters the runtime function.
    • 10_nested_loops, 06_math_intensive — parity by interleaved minimums; closed.

    Every number in this thread is from a named SHA, interleaved where deltas were small, with the instruments' own failures documented alongside the results.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions