You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
GC: arena-bytes trigger re-arm depends on which collector ran (tsc peak RSS ±28 MB knife-edge) #11736
The arena-bytes GC trigger re-arms at a point that depends on which collector discharged the pressure, not on the heap. The same heap re-arms about 50 MB apart depending on whether a budgeted non-moving minor or a precise-safepoint copying minor ran. On TypeScript (tscwork, N3) that turns a one-block difference in object placement into a +28 MB (+7%) peak RSS. The heap's live data is unchanged.
Evidence
Quiet dedicated host (EPYC 9275F, performance governor, boost off), release builds of main 9e29f59d43 and of #11713 (21ccb89264). #11713 changes no line in gc/policy.rs or arena/ (git diff 9e29f59d43 21ccb89264 -- crates/perry-runtime/src/gc/policy.rs crates/perry-runtime/src/arena is empty).
Live data is the same on both. At cycle 94: 28,784,416 B / 242,927 objects vs 28,784,688 B / 242,929. It matches at every full collection.
The memory is committed, not merely unpurged.MIMALLOC_SHOW_STATS=1 reports a committed peak of 317 vs 350 MiB, and MIMALLOC_PURGE_DELAY=0 still gives 384 vs 411 MB.
Main takes the high mode too when perturbed: one main run with GC tracing on peaked at 417 MB.
It disappears when both take the same path: with PERRY_GC_INCREMENTAL=0, peaks are 353–360 MB on both.
The two modes split at the first trigger after full collection 93 (cycle 94, about 5.8 s), and the split is consistent across 18 PERRY_GC_DIAG traces:
arena_total − next_base
collector
empty nursery blocks
arena trigger re-armed at
low mode (main, 9/9)
0 KiB
budgeted non-moving minor (site=budgeted_start kind=due)
copying minor at safepoint (site=safepoint kind=ArenaBytes)
62 stay reserved
~196 MB
In the high mode, the second allocation burst runs 3 consecutive arena minors with reservation going 159 → 183 → 191 → 208 → 222 MB before the old-gen full fires. That produces a peak of about 0.2 s.
Cause
gc_rebaseline_arena_trigger_after_collection (crates/perry-runtime/src/gc/policy.rs, around line 2785 at 9e29) re-arms with next_trigger = max(min(new_total + step, ceiling), new_total + floor), where new_total = arena_total_bytes().
What new_total counts: the empty nursery from-space blocks the copying minor keeps. It does not count the blocks the budgeted minor's general reclaim releases (arena/reset.rs).
Which collector catches the trigger:
The budgeted due-check is total >= next_arena_trigger_base() (around line 3752).
The safepoint arena path arms only when a block acquisition crosses the base.
So total == base is caught by the budgeted host poll, and one more block is caught by the copying minor.
Proposed fix
Make the re-arm independent of the collector. Derive it from committed bytes excluding empty, reusable nursery from-space, plus the nursery band and the step. Alternatively, have the copying minor drain empty from-space blocks above the nursery band, as the budgeted general reclaim already does.
Acceptance: the four-arm GC workload matrix (39 cells) plus tsc and Zod RSS, CPU and full counts. After the fix, main and #11713 should match in tsc RSS and fulls.
Related retained-capacity evidence from a different workload/platform; the causal link to this ticket's re-arm asymmetry is not established:
Measured 2026-10-01 on macOS arm64, Perry main d40ed1a47019bcae547819ebb5756bcd96ba1057 (v0.5.1655), with matching compiler and runtime archives. This snapshot already includes #11645. These are diagnostic probes; no compiler/runtime fix was applied and the original comparison remains unchanged.
The original cyclic workload peaks at 131.63 MiB RSS, three-run median. A separate full-GC census finds only approximately 5.16 MiB reachable tracked heap data while the cache is populated; after release and an explicit diagnostic collection it finds 1,584 bytes and no DocNodes. This demonstrates reclamation of this graph under explicit collection, not indefinite leak freedom.
The ordinary GC diagnostic run exits with 72 MiB arena capacity and reports no arena-right-size capacity return. VM maps of the unmodified workload show mimalloc Memory Tag 240 at approximately 117 MiB resident during churn and 124 MiB at final drain. The large reserved VM region has zero resident pages. Tag 240 includes arenas, allocator pages, Rust tables and scratch, not just live JS data; the views do not form an exact disjoint RSS ledger.
Reducing PERRY_GC_SCAVENGE_NURSERY_MB to 4 gives 78.52 MiB median peak RSS against 131.63 MiB default, with CPU increasing 10.64 → 11.90 s. Separate diagnostics show 36 versus 72 MiB exit arena capacity. The knob changes headroom/collection frequency, so it does not isolate allocator page purging.
#9709's idle right-sizing behavior already exists in this snapshot. This is a bounded actively churning/releasing workload, not the same idle soak or evidence that the existing idle fix regressed. These results support checking empty reusable nursery capacity and trigger re-arm behavior here, but no collector-mode comparison or re-arm trace has yet proven #11736's mechanism on this workload.
Compiler-side opportunity and complete unchanged reproducer: #11743.
Evidence in the project-comparison workspace: demo/results/stress/perry-profile/summary.md, profile-manifest.json, cpu-summary.json, experiment-summary.json, experiments.json, gc-diagnostics.stderr, nursery-4-gc-diagnostics.stderr, census.jsonl, vmmap-*.txt, and the original/typed-receiver makeTree disassemblies. These paths are local artifacts, not publicly hosted links.
Summary
The arena-bytes GC trigger re-arms at a point that depends on which collector discharged the pressure, not on the heap. The same heap re-arms about 50 MB apart depending on whether a budgeted non-moving minor or a precise-safepoint copying minor ran. On TypeScript (
tscwork, N3) that turns a one-block difference in object placement into a +28 MB (+7%) peak RSS. The heap's live data is unchanged.Evidence
Quiet dedicated host (EPYC 9275F, performance governor, boost off), release builds of main
9e29f59d43and of #11713 (21ccb89264). #11713 changes no line ingc/policy.rsorarena/(git diff 9e29f59d43 21ccb89264 -- crates/perry-runtime/src/gc/policy.rs crates/perry-runtime/src/arenais empty).MIMALLOC_SHOW_STATS=1reports a committed peak of 317 vs 350 MiB, andMIMALLOC_PURGE_DELAY=0still gives 384 vs 411 MB.PERRY_GC_INCREMENTAL=0, peaks are 353–360 MB on both.The two modes split at the first trigger after full collection 93 (cycle 94, about 5.8 s), and the split is consistent across 18
PERRY_GC_DIAGtraces:site=budgeted_start kind=due)[gc-general-reclaim] released=63)site=safepoint kind=ArenaBytes)In the high mode, the second allocation burst runs 3 consecutive arena minors with reservation going 159 → 183 → 191 → 208 → 222 MB before the old-gen full fires. That produces a peak of about 0.2 s.
Cause
gc_rebaseline_arena_trigger_after_collection(crates/perry-runtime/src/gc/policy.rs, around line 2785 at 9e29) re-arms withnext_trigger = max(min(new_total + step, ceiling), new_total + floor), wherenew_total = arena_total_bytes().new_totalcounts: the empty nursery from-space blocks the copying minor keeps. It does not count the blocks the budgeted minor's general reclaim releases (arena/reset.rs).total >= next_arena_trigger_base()(around line 3752).total == baseis caught by the budgeted host poll, and one more block is caught by the copying minor.Proposed fix
Make the re-arm independent of the collector. Derive it from committed bytes excluding empty, reusable nursery from-space, plus the nursery band and the step. Alternatively, have the copying minor drain empty from-space blocks above the nursery band, as the budgeted general reclaim already does.
Acceptance: the four-arm GC workload matrix (39 cells) plus tsc and Zod RSS, CPU and full counts. After the fix, main and #11713 should match in tsc RSS and fulls.