Skip to content

gc: the base ArenaBytes arm paces the nursery by a quantity that is 99 % old generation (#7909's young basis, not applied to the base arm) #9839

Description

@proggeramlug

Summary

gc_budgeted_due_trigger's base ArenaBytes arm tests arena_total_bytes()
— young generation + old generation + large objects — and schedules a
nursery minor, which can act on the young part only. Measured on the
compiled claude-code TUI, the young part is 0.9 % of the quantity the arm
tests at the median firing.

The sibling arm was already moved off exactly this basis. #7909 split
young_scavenge_cap_due() out of ArenaBytes precisely because
"the quantity this one tests is one a budgeted low-pause NON-MOVING cycle
cannot lower", and its doc comment predicts this data. The split was never
applied to the base arm, which still paces the nursery by the whole arena.

This is the basis half of the arm's design defect. #9831 / #9838 are the
step half, and they are complementary, not alternatives: #9838 fixes who
pulls the trigger down
(the tiny-parse pressure guard) and deliberately leaves
the arm's own re-arm alone, recording in a new comment that above
ceiling - floor the arithmetic re-arms at new_total + floor whatever step
says. That closes the step route by measurement (−10.8 % CPU for +22 % settled
footprint, the forbidden trade). It leaves the basis untouched.

Measured

Five PERRY_GC_DIAG=1 captures, two independently built binaries
(cc_main_0905, cc_int_0905), 3300- and 400-character streamed replies.
Per firing (not per trigger evaluation — the decision-input line describes a
different population), nursery occupancy is copied_bytes + promoted_bytes + freed_bytes from the firing's own [gc-copy-minor] ran line, against
pre_in_use from the [gc-step] line that same collection emits.

capture ArenaBytes firings median YOUNG share of the tested quantity old / large-object share
base 3300 i0 49 0.9 % 99.1 %
base 3300 i12 50 0.9 % 99.1 %
integ 3300 i0 47 0.9 % 99.1 %
integ 3300 i12 50 0.9 % 99.1 %
integ 400 42 1.5 % 98.5 %

Stable to the digit across two binaries and two workload sizes. Script:
secret-tests/cc-perf-campaign scratch basis_an.py.

The three shapes the arm actually fires in

Splitting the same captures by phase (at the first [gc-step] whose
pre_in_use reaches 90 MB, where the streamed turn's allocation ramp begins)
separates three mechanisms that the single name ArenaBytes hides:

shape what crossed where it occurs nursery at the firing
(a) request-lowered nothing — the trigger was lowered to "now" by the tiny-parse pressure guard after a small JSON.parse; in_use is flat across ~30 consecutive firings pre-turn, and the whole 400-char turn < 1 MB in 23 of 42
(b) crossed by the previous collection's promotion six MallocCount minors promoting ~3.2 MB each grow the old generation past a threshold no collection refreshed the 3300-char streaming turn, 8 of ~60 collections 856 bytes (promoted_bytes=216 freed_bytes=640), every time
(c) genuine whole-arena growth, quiet nursery 16 MB of old-generation or large-object growth with no collection in between does not occur in any cc capture —

Shape (a) is #9831/#9838's. Shape (b) is a finisher asymmetry
(gc_finish_malloc_trigger_collection never re-baselines the arena trigger)
and is being fixed separately. Shape (c) is this issue's, and it is the only
one where the arm's basis is the defect rather than a symptom: a nursery minor
on it frees nothing, because there is nothing young to free.

Why the other routes are closed

Four variants were built and measured; all are recorded in
secret-tests/cc-perf-campaign/HANDOFF_arenabytes_solution_space.md with the
branches (refuted/arena-*, head-stamped REFUTED — DO NOT LAND):

  • price the headroom by the step — −10.8 % CPU, +22 % settled footprint.
    The forbidden trade; now documented in-tree by fix(gc): price the tiny-parse pressure guard by the productivity backoff (#9831) #9838's comment.
  • yield to an arm that can act — measured inert: the arms each
    re-baseline when they fire, so they interleave rather than coincide. At the
    moment ArenaBytes is due, MallocCount has just re-baselined. 100 minors and
    a 50/50 arm split in both arms, identical to the digit. There is nothing to
    yield to.
  • route the residual to the budgeted old-gen arm — the routed arm is
    OldReclaim, which is evaluated first and has already declined at that
    moment. An arm that fires a full exactly when the full's own pacer says "not
    yet" is a second pacer with a worse constant. It also cannot clear its own
    cause: BudgetedGcRebaseline::OldReclaim never touches
    GC_NEXT_TRIGGER_BYTES, so it re-fires on every evaluation.
  • gate on young occupancy (the shape proposed below) — −21.7 % turn CPU at
    400 chars, −4.4 % at 3300, footprint and RSS improving at both
    , minors
    100 → 78 with total reclaimed up 3.5 %. It fails 26 tests in 7 modules,
    because arena pressure with a quiet nursery is then served by nothing.

That last number is why this is worth filing rather than dropping: the measured
win is real and the objection is a genuine hole, not a stale contract.

Proposed restatement

Arena pressure schedules, on the observing call, the collection that can act
on what the arm measures: a nursery minor when the young generation holds at
least a block, the escalating full when the pacing reading says so — and
otherwise nothing but a re-baseline, because a whole-arena total with a quiet
nursery is old-generation growth, which the old-generation arms pace.

Both old-generation quantities already have arms that can act on them
(OldReclaim's proportional band and arena_growth_full_escalation_due), and
both are evaluated before this one, so the declined case is handed on, not
dropped. The re-baseline on decline is what stops the once-per-block cadence the
routed variant hit.

What it would take

The 26 failing tests were read one by one (table in
secret-tests/cc-perf-campaign/DESIGN_arena_contract.md §1). No test's
invariant is "a minor ran on a quiet nursery."
Twenty are machinery
invariants — bounded stepping, phase parking, born-marked allocation,
drain-before-manual-gc, the atomic-finalize remark, trace shape, FFI safety,
root survival — for which make_arena_trigger_due() is a fixture: the
cheapest way to make the collector do something. Six assert Minor, and of
those only dirty_store_workload_reports_remembered_set_and_ordinary_pauses
needs a minor in substance (it verifies remembered-set telemetry, which only a
minor produces).

So the change is one fixture plus two tests:

  • make_arena_trigger_due() also guarantees ≥ BLOCK_SIZE of young occupancy,
    via the force_next_general_arena_alloc_slow idiom eight sibling
    runtime_roots tests already use — which is why those eight passed under the
    variant while the 26 did not.
  • two new tests in the gc: incremental old-gen work costs 14% on a program that never collects (asyncpipe, zero GC cycles) #7909 two-phase shape, so the decline is attributed
    rather than merely absent: the quiet-nursery decline re-baselines above total
    and starts nothing; the same heap with a block of young occupancy starts the
    minor.
  • Sabotage: remove the young gate → the first fails (the arm fires a minor
    on an empty nursery); remove the re-baseline → the second fails (the arm stays
    due, the once-per-block cadence).

Two contract statements move with it and belong in the changelog: the
js_gc_memory_pressure level-1 contract ("trigger lowered, collection deferred
to the next check") becomes "…if the nursery can be acted on" (level ≥ 2 sets
GC_OLD_RECLAIM_PENDING and is unaffected), and the tiny-parse channel's
cadence is unchanged.

Why an issue and not a PR

On cc this buys nothing today. Shape (c) does not occur in any capture, and
for shape (a) the young gate is identical in effect to #9838 — the guard
re-lowers the trigger at the next parse, so gating the arm merely moves the
firing to the next block rather than removing it. #9838 removes those
collections at the source, and it is the better fix for that shape. The young
gate's measured −21.7 % at 400 characters is the same prize #9838 claims;
only one of the two can have it, and they must not be measured as independent.

What this change buys is architectural: it closes the shape-(c) hole for
programs shaped like #5476 (4 M small allocations → 1.9 GB RSS when nothing
serves pressure), where a useless minor runs per 16 MB of large-object growth
today, and it makes the 26 fixtures honest about what they are testing.

Caveat on every 400-character figure above: they are figures on the current
JSON.parse
. Shape (a)'s input is one parse per SSE delta; a parser that
allocates differently moves the number in either direction. The mechanism — an
arm testing a quantity 99 % of which it cannot act on — does not depend on the
parser.

Activity

  1. proggeramlug commented on Sep 16, 2026

    @proggeramlug
    ContributorAuthor

    Second workload, same defect — plus a self-contained 20-line reproducer

    Confirming this on something much smaller than the cc TUI, measured on main 33690c563. The fixture is plain JS (gc3 from #10362: 300 000 allocated 8-node chains into a rolling ring of 40 000, no dependencies), so the arm can be studied without a compiled app.

    PERRY_GC_DIAG=1, per copying minor:

    minor survival copied promoted freed pause
    1–5 1000 ‰ 0 11–45 MB each 0 bytes 2–9 ms
    6–9 ~680 ‰ 14–30 MB 0–30 MB ~14 MB each 119–176 ms

    Triggers over the run: 10 ArenaBytes, 3 OldGenBytes, plus 4 fulls.

    The first five collections reclaim nothing at all — they are pure promotion, 115 MB of it, into an old generation that then needs the fulls to clear. That is this issue's shape (c) with a quiet-nursery twist: the arm keeps firing nursery minors while the quantity it tests grows because of what the previous minor promoted. Promote → arena grows → arm fires → promote more. On cc you measured the young share of the tested quantity at 0.9 %; here the loop is visible end to end in a single 20-line program.

    The other half: the cap ceiling

    NURSERY_CAP_SCALE_MAX = 4 (16 MB × 4) is reached on this workload, and the growth rule keys on influx (eden_live > cap/25) rather than on whether collections are reclaiming anything. Both facts point the same way: the nursery is below the workload's death window (the ring only drops a chain after 40 000 more are allocated), so everything is still reachable at collection time.

    Sweeping the base cap, holding allocation constant (instructions:u / peak RSS / max pause / collections):

    default 16 MB 32 MB 128 MB 256 MB
    gc3 (40k retained) 12.57 G / 225 MB / 174 ms / 9 5.77 G / 227 MB / 116 ms / 5 4.34 G / 214 MB / 204 ms / 2 3.19 G / 250 MB / 209 ms / 1
    w1000 (1k retained) 1.09 G / 65 MB / 8 ms / 11 1.10 G / 67 MB / 10 ms / 10 1.08 G / 141 MB / 20 ms / 2 1.14 G / 190 MB / 26 ms / 1
    oldyoung 1.53 G / 85 MB / 33 ms / 3 1.33 G / 83 MB / 35 ms / 2 0.53 G / 81 MB / — / 0 0.53 G / 82 MB / — / 0
    alloc-only 320 M / 47 MB / 3 ms / 7 361 M / 55 MB / 6 ms / 3 270 M / 94 MB / — / 0 270 M / 94 MB / — / 0

    Two things worth taking from that table:

    1. A bigger constant is not the answer. It is nearly free on gc3 (−54 % instructions at 32 MB for +2 MB RSS and a lower max pause) and costs w1000 three times the memory for nothing. Any fix has to be evidence-driven, not a new constant.
    2. The signal to drive it already exists per cycle. freed_bytes ≈ 0 with survival_permille ≈ 1000 says exactly "this collection reclaimed nothing because the nursery is smaller than this workload's death window". w1000 and alloc never show it; gc3 shows it five times in a row.

    I'm drafting a policy change along those lines (grow only on that signal, bounded by the existing tenured/2 term and a memory budget, plus this issue's basis fix so promoted bytes stop pacing the nursery). Happy to hand it over instead if this is already someone's — the issue reads unclaimed, and nothing is implemented yet.

    Reproducer, node/bun-identical, expected checksum=-606613590 live=-593353216 nodes=320000:

    var CHAINS = 40000, CHAIN_LEN = 8, TOTAL = 300000;
    function makeChain(seed) {
      var head = null;
      for (var i = 0; i < CHAIN_LEN; i++) head = { id: seed + i, payload: [seed, i, (seed ^ i) | 0, (seed + i * 3) | 0], next: head };
      return head;
    }
    var ring = new Array(CHAINS);
    for (var i = 0; i < CHAINS; i++) ring[i] = null;
    var checksum = 0;
    for (var n = 0; n < TOTAL; n++) { var c = makeChain(n); var tag = "n" + (n % 1024); checksum = (checksum + c.payload[n & 3] + tag.length) | 0; ring[n % CHAINS] = c; }
    var live = 0, nodes = 0;
    for (var i = 0; i < CHAINS; i++) { var cur = ring[i]; while (cur !== null) { live = (live + cur.id) | 0; nodes++; cur = cur.next; } }
    console.log("checksum=" + checksum + " live=" + live + " nodes=" + nodes);

    To see the diagnostics at all: perry compile strips the diagnostics feature via auto-optimize, so build the fixture with PERRY_NO_AUTO_OPTIMIZE=1 PERRY_RUNTIME_DIR=<tree>/target/release PERRY_LIB_DIR=<same> (without the DIR overrides it picks up a prebuilt archive and refuses as stale).

  2. proggeramlug commented on Sep 16, 2026

    @proggeramlug
    ContributorAuthor

    Correction: my attribution above was wrong — the base arm fired zero times on that fixture

    I claimed "10 ArenaBytes, 3 OldGenBytes" and "four of the nine minors fired below the nursery cap". Both are wrong, and the error is worth recording because the label invites it.

    trigger=ArenaBytes on the [gc-copy-minor] line does not identify the arm. The safepoint handler maps both BudgetedGcTrigger::ArenaBytes and BudgetedGcTrigger::YoungScavengeCap to GcTriggerKind::ArenaBytes (policy.rs:3589), and diag_sites::trigger_decision stamps the same string (policy.rs:3636). The [gc-trigger] line is what separates them. Re-reading my own capture:

    [gc-trigger] site=safepoint kind=ArenaBytes arena_total=14680064 next_base=134217728
                 from_space=11534256 nursery_cap=10938744
    

    arena_total (14.7 MB) is nowhere near next_base (134 MB); from_space (11.53 MB) is at 105 % of nursery_cap (10.94 MB). That is the cap arm, due on its own young basis. Across all 13 decisions on gc3, arena_total never reaches next_base: 0 base-arm firings, 10 cap-arm, 3 PromotedCohort fulls.

    Across six fixtures: 1 base-arm firing in 47 nursery-churn decisions (w5000), and that one is healthy — young was 84 % of the tested quantity and the minor freed 36.7 MB at 94 ‰ survival. It doubles as a positive control that the classifier can see base-arm firings, so the zeros are real.

    My "below the cap" claim was a cross-line pairing error: I took desired=1396703 from a [gc-tenuring] sweep-seed line emitted two minors later and paired it with minor 1's Eden. At minor 1 the cap was 10,938,744, not 22.3 MB — #8122's object census had re-denominated it down (mean_object_bytes 72 -> 47, scale 652 ‰). Every minor fired at 102–105 % of its cap. None below.

    What the fixture does show

    The numbers themselves reproduce exactly (minors 1–5 survival=1000 copied=0 freed=0, promoting 11.5/11.5/23.1/24.1/45.1 MB; minors 6–9 survival ~680, freed ~14.4 MB, pauses 121–185 ms). The mechanism is not pacing:

    • Minors 1–5 have copy_evacuation=0/0/0 with promotion=…/240299/11534256 — whole-block promotion. Eden's blocks go to the old generation without liveness examination, so survival=1000, freed=0 is definitional rather than an outcome.
    • The transition at minor 6 is exact: copied_bytes=30,720,672 = 40,000 chains × 768 B = the ring's entire retained set. The death window is 30.7 MB. No minor can reclaim anything until Eden exceeds it and takes the copying path.

    So this fixture is evidence for the cap-growth and promotion-path question, not for this issue's basis. It does not belong to #9839 and I withdraw it as such.

    On the basis fix itself

    For anyone picking this up: secret-tests/cc-perf-campaign/DESIGN_arena_contract.md §4 already works the solution space (design C), and records that variant 1 — the plain young gate — fails 26 tests in 7 modules, and that shape (c) "does not occur in any cc capture". It does not occur in any of my six either (0/47). Design C stays architecturally right, and its measurable value on the workloads either of us has is flat by construction. It would need a fixture that actually produces shape (c) (large-object or born-old growth with a quiet nursery, #5476's shape) before a gate table on it means anything.

    Apologies for the noise on the issue. The measurement stands; the attribution was mine and it was wrong.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions