Repository navigation
GC nursery pacing (#11645): accepted regressions to recover #11699
Description
Activity
- addedbugConfirmed defect or regressionConfirmed defect or regressionperformanceRuntime, compile-time, build-size, or memory performanceRuntime, compile-time, build-size, or memory performance
on Sep 30, 2026 Re-measured after #11676 merged. Main
034b1ea38against #11645 at1412ce8f2, with the same method as the issue body: THP off, the bench lock, and median of 11. Raw data is in/root/claude-gc-nursery-pacing-bench/m8.txton perrymaster.row Δ instr (was) Δ instr now Δ RSS (was) Δ RSS now minors/fulls, main → PR binary-trees n=3 +746% +497.7% (24.96 M → 149.18 M) +63.8% +60.2% 0/0 → 1/0 binary-trees n=6 +649% +432.9% +59.0% +54.6% 0/0 → 1/0 binary-trees n=10 +553% +368.7% +54.7% +49.1% 0/0 → 1/0 gc_ratchet 02 +17.3% +19.8% (257.56 M → 308.65 M) −11.6% −12.0% 1/1 → 2/1 gc_ratchet 12 −5.7% −5.1% +2.6% +2.5% 5/1 → 6/1 moment −20.6% −20.4% +3.1% +13.8% (37.4 → 42.5 MB) 0/33 → 43/0 (main had 1/33) qs parse_nested +0.36% +0.75% +1.2% −2.4% 11/0 → 14/0 alloc loop +0.57% +0.57% −20.0% −19.3% 67/0 → 203/0 - binary-trees shrank as predicted. The instruction cost fell by about a third, to roughly the +474/+412/+351% measured when perf(gc): survival-aware nursery pacing, 4 MB floor on the existing influx ladder (includes #11612) #11645 was stacked on perf(gc): cut the copying minor's per-live-object trace cost (straight-line object scan, hoisted per-object facts, exact memos) #11676. The RSS cost is unchanged, which fits: it comes from the minor's tracking lists, not from the trace.
- 02 did not shrink. perf(gc): cut the copying minor's per-live-object trace cost (straight-line object scan, hoisted per-object facts, exact memos) #11676 made main's own collections cheaper too (277 M → 258 M), so the second minor's relative cost stayed about the same. The earlier stacked "−2.8%" reading for 02 came from a different comparison and does not hold on main.
- moment RSS got worse: +3.1% → +13.8%. Main tightened to 37.2–38.5 MB, now with 0 minors where it previously ran 1. The PR went from 40.8 MB to 42.5 MB (range 42.2–42.9). This is now the largest unexplained item here. The bisect suggested in the body should cover both arms over
10ece9958..034b1ea38, including perf(gc): cut the copying minor's per-live-object trace cost (straight-line object scan, hoisted per-object facts, exact memos) #11676. - qs parse went from +0.36% to +0.75%. That is still small, but it doubled.
Correctness on the rebased head is clean:
cargo test -p perry-runtimepasses.- The seeded
PROTECT_FROMSPACEsweeps pass on dotenv, moment, qs and binary-trees: 200/200 on the previous base and 50/50 on the final head. - The gc/array/regex gap A/B matches main.
#11645 is now ready at
356e189a3, on main5fbc2c3b0. The numbers above were not re-taken after that last rebase: main's P4 shape-only GC flip is in between and could move them.Owner decision 2026-09-30: #11645 lands with moment peak RSS at +13.8% (37.4 → 42.5 MB, main 034b1ea vs PR, THP off, median of 11). That is above the +3.1% accepted earlier. The change comes from main itself: main now runs 0 minors on moment, where it used to run 1. Recovering it is tracked here, together with gc_ratchet 02 at +19.8% instructions. Next step: bisect main between 7fa094c and 034b1ea for moment's minor-count change.
Additional nursery/headroom tradeoff on a cyclic graph workload, measured after #11645 landed:
Measured 2026-10-01 on macOS arm64, Perry main
d40ed1a47019bcae547819ebb5756bcd96ba1057(v0.5.1655), with matching compiler and runtime archives. This snapshot already includes #11645. These are diagnostic probes; no compiler/runtime fix was applied and the original comparison remains unchanged.Arm Median CPU seconds Median peak RSS MiB Original/default policy 10.64 131.63 Original, PERRY_GC_SCAVENGE_NURSERY_MB=811.42 99.42 Original, PERRY_GC_SCAVENGE_NURSERY_MB=411.90 78.52 Typed child-array local/default 4.90 131.55 Typed local + nursery base 4 MiB 6.03 79.58 Three sequential repeats per arm; CPU = child user + system time from wait4, RSS = peak resident memory. Every run matched the original full stdout and exit status, including parent identity and closure/cache checks. Default, nursery-8, nursery-4 and typed-local arms rotated order across three rounds; the combined arm ran afterward. Same constants, node count and 1 GiB limit.
Reducing the base alone costs approximately 12% CPU while reducing RSS approximately 40%. Separate GC diagnostic captures report 981 copying minors at base 4 versus 274 at default, and exit arena capacity 36 MiB versus 72 MiB. Diagnostic timing/counts are not substituted for uninstrumented CPU/RSS. The combined arm is 43.3% lower CPU and 39.5% lower RSS versus fresh default; its CPU win comes from a separate compiler-call-lowering probe.
The explicit base override is not equivalent to #11645's 4 MB starting influx-ladder floor: the base can be dynamically scaled and is not the total/effective nursery size. This is residual workload evidence after the adaptation landed, not a pre-#11645 comparison or a recommendation to lower defaults globally. It does not establish which additional policy change would recover the CPU cost.
Compiler-side opportunity and complete unchanged reproducer: #11743.
Evidence in the project-comparison workspace:
demo/results/stress/perry-profile/summary.md,profile-manifest.json,cpu-summary.json,experiment-summary.json,experiments.json,gc-diagnostics.stderr,nursery-4-gc-diagnostics.stderr,census.jsonl,vmmap-*.txt, and the original/typed-receivermakeTreedisassemblies. These paths are local artifacts, not publicly hosted links.
The owner accepted these regressions on 2026-09-30 when landing #11645 ("land 11645 now, but let's have an issue for the small regression with as many details as we can right now"). This issue tracks getting them back. Part of #11549.
#11645 (survival-aware nursery pacing) starts the scavenge nursery ladder at 4 MB instead of 16 MB. That wins big on package workloads: dotenv −30% instructions, and validator, date-fns, jwt and uuid −20–30% peak RSS. It costs small or fully-live programs one or more extra minors. The rows below are the ones that got worse.
Numbers (main vs PR)
The rows in bold are the accepted regressions.
Method
--release(-p perry -p perry-runtime-static -p perry-stdlib-static), compiled withPERRY_NO_AUTO_OPTIMIZE=1, and run withTZ=UTC.7fa094cb4. The PR was0cce8784f, i.e. perf(gc): survival-aware nursery pacing, 4 MB floor on the existing influx ladder (includes #11612) #11645's five commits (HOLD (RSS trade): perf(regex): stop charging operation scratch to GC pressure; grow lent registers (bounded) #11612 ×2, pacing, root-holders classify, changelog) on that main. perf(gc): survival-aware nursery pacing, 4 MB floor on the existing influx ladder (includes #11612) #11645 was later rebased onto6a5090751as427728113; the numbers were not re-taken on that base.MIMALLOC_ALLOW_THP=0, i.e. 4 KiB accounting). The host has THP inmadvisemode, and THP-on RSS changes from run to run by several %; see perf(gc): survival-aware nursery pacing, 4 MB floor on the existing influx ladder (includes #11612) #11645's body.perf stat -e instructions:u. Package and loop rows are per iteration, from a two-N differential (median of 3 at n1, median of 11 at n2). binary-trees and gc_ratchet rows are totals for a fixed size./usr/bin/time %Mon the binary, median of 11, arms interleaved. moment, 12 and 02 were re-run at 21 reps.PERRY_GC_DIAG=1./tmp/perry-bench-lock.d./root/claude-gc-nursery-pacing-bench/m6.txt: instructions and RSS, all rows./root/claude-gc-nursery-pacing-bench/m7.txt: the 21-rep RSS re-run of moment, 12 and 02./root/claude-gc-nursery-pacing-bench/main2/and/root/claude-gc-nursery-pacing-bench/pr2/./root/claude-gc-nursery-pacing-{meas2,meas3,compile}.sh, tabulated bybench/tab.pyandbench/tab3.py.benchmarks/packages/*(dotenv 18.0.1, qs 6.16.0, moment 2.31.0, …) andbenchmarks/gc_ratchet/probes/{01,02,12}. binary-trees, alloc and retain are the.tsfiles in the bench dir.Per-row mechanisms
visit_gc_layout_slot_descriptors,scan_object_fields,classify_arena,visit_slot_with_weak_fact), plus ~33 M of non-trace fixed setup. Main pays the same per-object cost from n≥20 on, where n=20 is −4.0% here. The extra RSS comes from the minor'sworklistandmoved_headersdoubling from a zero estimate. A never-reallocating header list was tried onwip/11549-header-list-experiment: it cut about 3 MB, but traded RSS elsewhere under THP (binary-trees 10→40 +7%, gc_ratchet 01 +3.6%), so it was not landed.gc()full then sweeps a young generation still full of garbage. The old perf(gc): survival-aware nursery pacing, 4 MB floor on the existing influx ladder (includes #11612) #11645 table showed +4.3%, but that compared main without perf(gc): the first collection's barrier-arming walk skips wholly-nursery blocks #11668 against the PR with it. With perf(gc): the first collection's barrier-arming walk skips wholly-nursery blocks #11668 on main (d7df6e7), this is the second minor's true cost.Unexplained: moment's RSS shift
On
10ece9958(the previous #11645 measurement), moment's THP-off RSS was main 40.8 MB against PR 35.3 MB (−13.5%). On7fa094cb4, main is 39.6 MB (range 37.7–45.2) and the PR is 40.8 MB (range 40.0–41.2). Main barely moved; the PR's median rose ~5.5 MB. Minor and full counts are unchanged (1/33 against 43/0), and the −20.6% instruction win still holds.Nothing has been bisected. The candidates are the GC- or rooting-touching commits in
10ece9958..7fa094cb4:Func.prototype.x = <call>. moment does this pervasively at startup, and its peak is set in the first ~0.2 s by the startup cohort's copy.The first step is to bisect the PR arm over this range with
MIMALLOC_ALLOW_THP=0, looking at moment's peak RSS.Expected help from #11676 (per-object trace cost)
#11676 cuts the per-object trace ~35%. With #11645 stacked on it earlier, binary-trees came out at +474/+412/+351% (n=3/6/10) instead of +746/+649/+553%, and 02 at −2.8% instead of +17.3%. So once #11676 lands, 02 should drop off this list, and binary-trees should roughly halve.
Candidate next steps
Prune the fixed per-minor work that qs's minors spend time in. Shares of qs minor time: closure box-capture prune 16%, layout-owner prune/sort 12%, remembered-set rebuild 12%. Making these incremental or skipping them when nothing changed helps qs, the alloc loop and every high-minor-count row.
Cut the first minor's fixed setup (the ~33 M of non-trace work in binary-trees' single minor).
A survival-aware first-minor policy. At the first minor, survival alone does not separate binary-trees' fully-live tree from a package's startup cohort, which is 10–31% alive (0.4–1.3 MB). Measured alternatives:
A policy would need a signal beyond first-minor survival, e.g. live bytes relative to allocated bytes, or deferring the promotion decision.