Skip to content

GC nursery pacing (#11645): accepted regressions to recover #11699

Description

@proggeramlug

The owner accepted these regressions on 2026-09-30 when landing #11645 ("land 11645 now, but let's have an issue for the small regression with as many details as we can right now"). This issue tracks getting them back. Part of #11549.

#11645 (survival-aware nursery pacing) starts the scavenge nursery ladder at 4 MB instead of 16 MB. That wins big on package workloads: dotenv −30% instructions, and validator, date-fns, jwt and uuid −20–30% peak RSS. It costs small or fully-live programs one or more extra minors. The rows below are the ones that got worse.

Numbers (main vs PR)

row Δ instructions Δ peak RSS minors/fulls, main → PR
dotenv_parse −30.5% −20.3% 0/7 → 61/0
moment_parse_format −20.6% +3.1% (21 reps; +5.4% at 11 reps) 1/33 → 43/0
validator_batch −3.7% −26.0% 46/0 → 211/0
date-fns_format_add −1.0% −30.0% 17/0 → 71/0
qs_parse_nested +0.36% +1.2% (ranges overlap) 11/0 → 14/0
qs_stringify_nested −0.3% +0.6% 24/10 → 24/10
jsonwebtoken_decode −3.6% −22.0% 2/0 → 17/0
uuid_v4 −0.1% −21.0% 1/0 → 6/0
alloc loop +0.57% −20.0% 67/0 → 203/0
binary-trees n=3 +746% (24.95 M → 211.09 M) +63.8% (15.5 → 25.3 MB) 0/0 → 1/0
binary-trees n=6 +649% (28.69 M → 214.83 M) +59.0% 0/0 → 1/0
binary-trees n=10 +553% (33.68 M → 219.81 M) +54.7% 0/0 → 1/0
binary-trees n=20 −4.0% −8.9% 1/0 → 2/0
gc_ratchet 01_nursery_churn −4.5% −25.5% 1/1 → 4/1
gc_ratchet 02_survivor_promotion +17.3% (276.8 M → 324.6 M) −11.6% 1/1 → 2/1
gc_ratchet 12_large_live_set −5.7% +2.6% (105.6 → 108.4 MB) 5/1 → 6/1

The rows in bold are the accepted regressions.

Method

Per-row mechanisms

  • binary-trees n=3/6/10. The program builds a ~5 MB tree that stays fully live and exits after 7–9 MB of allocation. Main never collects it. The PR runs one minor at the 4 MB floor, and that minor promotes the whole tree in place: about +186 M instructions. Earlier profiling put this minor's cost at ~120 M of trace, which is the collector's general per-object visit/classify cost (visit_gc_layout_slot_descriptors, scan_object_fields, classify_arena, visit_slot_with_weak_fact), plus ~33 M of non-trace fixed setup. Main pays the same per-object cost from n≥20 on, where n=20 is −4.0% here. The extra RSS comes from the minor's worklist and moved_headers doubling from a zero estimate. A never-reallocating header list was tried on wip/11549-header-list-experiment: it cut about 3 MB, but traded RSS elsewhere under THP (binary-trees 10→40 +7%, gc_ratchet 01 +3.6%), so it was not landed.
  • gc_ratchet 02 (+17.3%). The PR runs 2 minors where main runs 1. Minor 2 re-copies minor 1's surviving cohort, and the final gc() full then sweeps a young generation still full of garbage. The old perf(gc): survival-aware nursery pacing, 4 MB floor on the existing influx ladder (includes #11612) #11645 table showed +4.3%, but that compared main without perf(gc): the first collection's barrier-arming walk skips wholly-nursery blocks #11668 against the PR with it. With perf(gc): the first collection's barrier-arming walk skips wholly-nursery blocks #11668 on main (d7df6e7), this is the second minor's true cost.
  • gc_ratchet 12 (+2.6% RSS). One extra early minor (6 vs 5) promotes part of the large live set earlier than main does.
  • alloc loop (+0.57%). 203 minors against 67. Each minor has a fixed cost of ~100 k instructions (perf(gc): cut a copying minor's fixed cost (intern young log, skip empty array-tail tables, young-only prunes) #11634 line of work). The 136 extra minors × ~106 k account for the whole delta. To break even, the fixed cost would have to fall to ~35 k per minor.
  • qs_parse_nested (+0.36%). 14 minors against 11. See the qs prunes below.

Unexplained: moment's RSS shift

On 10ece9958 (the previous #11645 measurement), moment's THP-off RSS was main 40.8 MB against PR 35.3 MB (−13.5%). On 7fa094cb4, main is 39.6 MB (range 37.7–45.2) and the PR is 40.8 MB (range 40.0–41.2). Main barely moved; the PR's median rose ~5.5 MB. Minor and full counts are unchanged (1/33 against 43/0), and the −20.6% instruction win still holds.

Nothing has been bisected. The candidates are the GC- or rooting-touching commits in 10ece9958..7fa094cb4:

The first step is to bisect the PR arm over this range with MIMALLOC_ALLOW_THP=0, looking at moment's peak RSS.

Expected help from #11676 (per-object trace cost)

#11676 cuts the per-object trace ~35%. With #11645 stacked on it earlier, binary-trees came out at +474/+412/+351% (n=3/6/10) instead of +746/+649/+553%, and 02 at −2.8% instead of +17.3%. So once #11676 lands, 02 should drop off this list, and binary-trees should roughly halve.

Candidate next steps

  1. Prune the fixed per-minor work that qs's minors spend time in. Shares of qs minor time: closure box-capture prune 16%, layout-owner prune/sort 12%, remembered-set rebuild 12%. Making these incremental or skipping them when nothing changed helps qs, the alloc loop and every high-minor-count row.

  2. Cut the first minor's fixed setup (the ~33 M of non-trace work in binary-trees' single minor).

  3. A survival-aware first-minor policy. At the first minor, survival alone does not separate binary-trees' fully-live tree from a package's startup cohort, which is 10–31% alive (0.4–1.3 MB). Measured alternatives:

    • Powering on at the 16 MB base fixes binary-trees and 02 but costs dotenv +22.6% and moment +25.1% RSS.
    • An 8 MB start is a size-specific fit that costs 02 +20.6%.

    A policy would need a signal beyond first-minor survival, e.g. live bytes relative to allocated bytes, or deferring the promotion decision.

Activity

  1. added
    bugConfirmed defect or regression
    performanceRuntime, compile-time, build-size, or memory performance
    on Sep 30, 2026
  2. proggeramlug commented on Sep 30, 2026

    @proggeramlug
    ContributorAuthor

    Re-measured after #11676 merged. Main 034b1ea38 against #11645 at 1412ce8f2, with the same method as the issue body: THP off, the bench lock, and median of 11. Raw data is in /root/claude-gc-nursery-pacing-bench/m8.txt on perrymaster.

    row Δ instr (was) Δ instr now Δ RSS (was) Δ RSS now minors/fulls, main → PR
    binary-trees n=3 +746% +497.7% (24.96 M → 149.18 M) +63.8% +60.2% 0/0 → 1/0
    binary-trees n=6 +649% +432.9% +59.0% +54.6% 0/0 → 1/0
    binary-trees n=10 +553% +368.7% +54.7% +49.1% 0/0 → 1/0
    gc_ratchet 02 +17.3% +19.8% (257.56 M → 308.65 M) −11.6% −12.0% 1/1 → 2/1
    gc_ratchet 12 −5.7% −5.1% +2.6% +2.5% 5/1 → 6/1
    moment −20.6% −20.4% +3.1% +13.8% (37.4 → 42.5 MB) 0/33 → 43/0 (main had 1/33)
    qs parse_nested +0.36% +0.75% +1.2% −2.4% 11/0 → 14/0
    alloc loop +0.57% +0.57% −20.0% −19.3% 67/0 → 203/0

    Correctness on the rebased head is clean:

    • cargo test -p perry-runtime passes.
    • The seeded PROTECT_FROMSPACE sweeps pass on dotenv, moment, qs and binary-trees: 200/200 on the previous base and 50/50 on the final head.
    • The gc/array/regex gap A/B matches main.

    #11645 is now ready at 356e189a3, on main 5fbc2c3b0. The numbers above were not re-taken after that last rebase: main's P4 shape-only GC flip is in between and could move them.

  3. proggeramlug commented on Sep 30, 2026

    @proggeramlug
    ContributorAuthor

    Owner decision 2026-09-30: #11645 lands with moment peak RSS at +13.8% (37.4 → 42.5 MB, main 034b1ea vs PR, THP off, median of 11). That is above the +3.1% accepted earlier. The change comes from main itself: main now runs 0 minors on moment, where it used to run 1. Recovering it is tracked here, together with gc_ratchet 02 at +19.8% instructions. Next step: bisect main between 7fa094c and 034b1ea for moment's minor-count change.

  4. proggeramlug commented on Oct 2, 2026

    @proggeramlug
    ContributorAuthor

    Additional nursery/headroom tradeoff on a cyclic graph workload, measured after #11645 landed:

    Measured 2026-10-01 on macOS arm64, Perry main d40ed1a47019bcae547819ebb5756bcd96ba1057 (v0.5.1655), with matching compiler and runtime archives. This snapshot already includes #11645. These are diagnostic probes; no compiler/runtime fix was applied and the original comparison remains unchanged.

    Arm Median CPU seconds Median peak RSS MiB
    Original/default policy 10.64 131.63
    Original, PERRY_GC_SCAVENGE_NURSERY_MB=8 11.42 99.42
    Original, PERRY_GC_SCAVENGE_NURSERY_MB=4 11.90 78.52
    Typed child-array local/default 4.90 131.55
    Typed local + nursery base 4 MiB 6.03 79.58

    Three sequential repeats per arm; CPU = child user + system time from wait4, RSS = peak resident memory. Every run matched the original full stdout and exit status, including parent identity and closure/cache checks. Default, nursery-8, nursery-4 and typed-local arms rotated order across three rounds; the combined arm ran afterward. Same constants, node count and 1 GiB limit.

    Reducing the base alone costs approximately 12% CPU while reducing RSS approximately 40%. Separate GC diagnostic captures report 981 copying minors at base 4 versus 274 at default, and exit arena capacity 36 MiB versus 72 MiB. Diagnostic timing/counts are not substituted for uninstrumented CPU/RSS. The combined arm is 43.3% lower CPU and 39.5% lower RSS versus fresh default; its CPU win comes from a separate compiler-call-lowering probe.

    The explicit base override is not equivalent to #11645's 4 MB starting influx-ladder floor: the base can be dynamically scaled and is not the total/effective nursery size. This is residual workload evidence after the adaptation landed, not a pre-#11645 comparison or a recommendation to lower defaults globally. It does not establish which additional policy change would recover the CPU cost.

    Compiler-side opportunity and complete unchanged reproducer: #11743.

    Evidence in the project-comparison workspace: demo/results/stress/perry-profile/summary.md, profile-manifest.json, cpu-summary.json, experiment-summary.json, experiments.json, gc-diagnostics.stderr, nursery-4-gc-diagnostics.stderr, census.jsonl, vmmap-*.txt, and the original/typed-receiver makeTree disassemblies. These paths are local artifacts, not publicly hosted links.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugConfirmed defect or regressionperformanceRuntime, compile-time, build-size, or memory performance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions