Skip to content

perf(gc): 88% of a real TypeScript transpile is GC — 244 full collections per transpileModule, because #7937's absolute old-reclaim arm re-arms after every collection (a ~42 MiB live set parked just under the 48 MiB threshold) #10928

Description

@proggeramlug

Summary

On a real TypeScript workload, ~88% of perry's instructions are the garbage collector, and
the cause is a pacing arm: 244 full collections in one ts.transpileModule call, each
tracing ~182 MiB of live old-gen to reclaim ~7.9 MiB. The arm that fires is the #7937
absolute first-crossing arm of old_reclaim_pressure_due. It is meant to fire once, when old-gen
first crosses 48 MiB, but it re-arms after every collection, because the post-reclaim
baseline (~42 MiB) never reaches the threshold. The result is the fixed-bytes cadence that
#7592's own comment says makes total major-GC work quadratic in the live set.

The pacing code is byte-identical between the measured build (v0.5.1631, f88acdaa1) and current
main (0fa391529).

The workload and the gap

/root/fr/work/tscwork.ts on perrymaster (lane 14's): builds a 1,200-line TypeScript source and
calls ts.transpileModule on it 3 times (typescript 5.8.2). Complete runs, both output
typescript 5.8.2 290280, identical to node.

instructions:u wall
node v26.5.1, min of 3 (2% spread) 7.13 G ~1.3 s user
perry, lane 14's complete run (perf stat, stripped binary) 1,599.70 G 418 s
perry, this attribution run (826,329 samples × period 2,000,003) 1,652.66 G 498 s under perf record

perry / node ≈ 224–232×. The two perry counts come from independent binaries and methods and
agree within 3%. This is a different workload from the 9.7× reported in September, which was
tsc --noEmit type-checking (scanner/parser over lib.*.d.ts), timed by wall clock on macOS
arm64.

Where the instructions go

perf record -e instructions:u -c 2000003 on a PERRY_KEEP_SYMBOLS=1 +
PERRY_NO_AUTO_OPTIMIZE=1 build, self time, bucketed by the demangled symbol's module path.
Relative standard error per bucket ≈ 1/√samples.

subsystem samples ≈ instr share ±SE
GC / allocation (perry_runtime::gc::*, arena, mimalloc) 706,719 1,413.4 G 85.53% 0.1%
+ weakref::is_weak_target_trace_slot (GC tracing, placed under "other" by the first pass) 23,860 47.7 G 2.89% 0.6%
object model (perry_runtime::object::*, closures, proxy, symbols) 60,285 120.6 G 7.30% 0.4%
other runtime (excluding the weakref row) 10,900 21.8 G 1.32% 1.0%
unclassified (quicksort<usize> 0.62%, SipHash ~0.2% — no call graph was recorded) 7,882 15.8 G 0.95% 1.1%
arrays 7,011 14.0 G 0.85% 1.2%
strings 6,066 12.1 G 0.73% 1.3%
compiled code (perry_fn_*, perry_method_*, closures) 2,542 5.1 G 0.31% 2.0%
maps / sets, regex, numbers, exceptions 1,064 2.1 G 0.13% —

GC ≈ 88.4% of the run. perry's own compiled code — 5.1 G — is less than node's entire
run (7.13 G). A zero-cost object model would leave perry at ~215× node.

Top 15 symbols: 14 are GC and the 15th is the weakref trace slot. IncrementalSweepState::step
19.7%, ValidPointerSet::contains 11.6%, trace_heap_rewrite_slots 8.1% + 4.0%,
visit_gc_rewrite_slots 6.1%, visit_gc_layout_slot_descriptors 5.9%,
remember_evacuated_old_to_young_slot 4.8%, remembered_child_needs_tracking 3.2%,
ValidPointerSetBuilder::step 3.2%, OldToYoungRememberedRebuildState::step 2.9%,
per_object_slot_mask 2.6%, … Two notes so nobody chases the wrong thing: gc::verify is the
module that holds the production forwarding rewrite and remembered-set rebuild, not a debug
verifier left on; and ValidPointerSet is already a per-block object-start bitmap with a fence
binary search. It is hot because it is called for every traced pointer, not because it is slow.

Why: the collection count

PERRY_GC_DIAG=1 (diagnostic only, no behaviour change) on one iteration:

event count
[gc-full] full collections 244
[gc] collections, all kinds 256
triggers site=alloc_point kind=OldReclaim 242
triggers safepoint kind=ArenaBytes 14
triggers budgeted_start due 12
[gc-copy-minor] 26

Every one of the 242 OldReclaim triggers reports the same old_threshold=50331648 (48 MiB). At the
trigger:

per full collection (median of 242)
watched quantity (old_reclaimable) at trigger 47.6 MiB
old_baseline (post-reclaim value) ~42 MiB
watched quantity reclaimed 7.9 MiB
live old gen the collection has to trace (old_in_use − old_reclaimable) 182.2 MiB
old_free — free old-gen space at the trigger 182.2 MiB
traced : reclaimed 23 : 1
GC instructions per full collection (≈ 88.4% × 550.9 G per iteration ÷ 244) ~2.0 G

It is not memory pressure: 182 MiB of old-gen is free every time.

The arm, from source (gc/policy.rs:1881, unchanged on main)

pub(super) fn old_reclaim_pressure_due(old_in_use: usize, baseline: usize) -> bool {
    let threshold = gc_old_gen_reclaim_threshold_dyn_bytes();          // 48 MiB
    let crossed_absolute_threshold = old_in_use >= threshold
        && baseline < threshold
        && !GC_MAJOR_PACING_RETAINING.with(|c| c.get());
    crossed_absolute_threshold
        || old_in_use.saturating_sub(baseline) >= gc_old_reclaim_growth_band_bytes(baseline)
}

With the measured values — old_in_use ≈ 47.6–50 MiB ≥ 48 MiB, baseline ≈ 42 MiB < 48 MiB, and
RETAINING false because tsc's young generation is dying — the absolute arm is true every
time
, so the proportional arm is never consulted. The #7937 comment on this very arm says "it is
the arm that actually fires"
, and exempts it only while the heap is RETAINING. After each full
collection the baseline resets to the post-reclaim value (~42 MiB). That is still below 48 MiB, so
the "first crossing" is re-armed, and the next ~7.9 MiB of watched growth crosses it again.

That is the shape #7592's comment on gc_old_reclaim_growth_band_bytes describes and was written
to remove: "each full reclaim costs O(live), so a fixed-bytes cadence makes total major-GC work
quadratic in the live set."
Here the fixed cadence arrives through the absolute arm instead of the
band.

One step further, and marked as inference rather than measurement. Removing the absolute arm
alone would not make the pacing proportional to this heap. The proportional arm's baseline is the
same ~42 MiB watched quantity, so its band is max(32 MiB floor, 42/2) = the 32 MiB floor: a full
collection per ~32 MiB of watched growth instead of per ~7.9 MiB, roughly 4× fewer. The ~182 MiB
live set never enters either side of the comparison. Whether the watched quantity
(old_gen_reclaimable_pressure_bytes) is the right thing to pace on is the deeper question, and it
is for the GC lane.

Relation to existing issues

Diagnosis only; nothing built or proposed as a patch. Artefacts on perrymaster:
/root/matrix/tsc_attr.data (perf record), /root/matrix/tsc_gcdiag.err (GC diag),
/root/matrix/attribute.py.

Activity

  1. proggeramlug commented on Sep 21, 2026

    @proggeramlug
    ContributorAuthor

    Correction: the live set is ~42 MiB, not ~182 MiB. The 182 MiB was the free list.

    I mislabelled a field in the body above. The pacing diagnosis stands. The traced-to-reclaimed
    ratio does not.

    From gc/policy.rs:1832:

    fn old_gen_reclaimable_pressure_bytes() -> usize {
        crate::arena::old_gen_in_use_bytes().saturating_sub(super::old_free_bytes())
    }

    So in the [gc-trigger] line, old_reclaimable is the occupied old gen (live plus garbage
    since the last collection), and old_in_use is the arena's whole old-gen footprint (occupied plus
    free list). I took old_in_use − old_reclaimable to be live data. It is identically old_free: the
    equality old_in_use == old_free + old_reclaimable holds exactly in all 242 OldReclaim
    triggers
    . I even printed both medians in the body as 182.2 MiB and did not notice they were the
    same number.

    Corrected, median of 242 full collections:

    was is
    live old gen each full collection traces (the post-reclaim old_baseline) 182.2 MiB ~42 MiB
    occupied old gen at the trigger (old_reclaimable) 47.6 MiB 47.6 MiB
    reclaimed per full collection 7.9 MiB 7.9 MiB
    traced : reclaimed 23 : 1 ~5 : 1
    free-list space in the old-gen arena (old_free) 182.2 MiB 182.2 MiB — free, not live

    The correction tightens the pacing explanation rather than weakening it. The live set (~42 MiB)
    sits just under the fixed 48 MiB threshold, so the headroom before the next "first crossing"
    is only ~6–8 MiB. That is why a full collection fires on every ~7.9 MiB of growth: a live set parked
    just under a fixed threshold is exactly the case a threshold-only trigger handles worst.

    Also struck from the body:

    • "The ~182 MiB live set never enters either side of the comparison." Wrong. The live set is
      the ~42 MiB baseline, and it enters as baseline. The inference beside it still holds: with a
      ~42 MiB baseline the proportional band is its 32 MiB floor, so removing the absolute arm alone
      gives roughly one full collection per ~32 MiB instead of per ~7.9 MiB, about 4× fewer.

    Unaffected: 244 full collections per transpileModule, 242 of 256 triggers OldReclaim, the arm
    that fires, ~2.0 G instructions per full collection, GC ≈ 88.4% of the run, ~224–232× node.

    With ~42 MiB of live data at ~2.0 G instructions per full collection, each collection costs ~48
    instructions per live byte. That is high for a mark over 42 MiB, and it matches the top symbol
    being IncrementalSweepState::step (19.7%, sweep, not mark): sweep cost scales with the
    arena's extent — occupied plus the 182 MiB free list — not with the live set. That is a
    second, separate lever, and I'm marking it as inference until measured.

  2. changed the title [-]perf(gc): 88% of a real TypeScript transpile is GC — 244 full collections per transpileModule, because #7937's absolute old-reclaim arm re-arms after every collection (traces 182 MiB to reclaim 7.9 MiB, 23:1)[/-] [+]perf(gc): 88% of a real TypeScript transpile is GC — 244 full collections per transpileModule, because #7937's absolute old-reclaim arm re-arms after every collection (a ~42 MiB live set parked just under the 48 MiB threshold)[/+] on Sep 21, 2026
  3. proggeramlug commented on Sep 21, 2026

    @proggeramlug
    ContributorAuthor

    Correction: old_in_use − old_reclaimable is old_free — the "182 MiB live traced" and "23:1" rows are not supported

    Source-derived, posted before my own diag run, because a composition measurement is being built
    against the 182 MiB figure right now.
    Pending empirical confirmation (below); I will post the
    confirmation or the retraction.

    The algebra

    old_reclaimable is defined as old_in_use − old_free — gc/policy.rs:1832:

    /// Old-gen pressure the reclaim arms act on: block-offset in-use minus the
    /// swept holes the free list can already hand back (#7437).
    pub(super) fn old_gen_reclaimable_pressure_bytes() -> usize {
        crate::arena::old_gen_in_use_bytes().saturating_sub(super::old_free_bytes())
    }

    and diag_sites.rs:46-57 prints old_in_use, old_free and old_reclaimable from exactly those
    definitions. So, as an identity:

    old_in_use − old_reclaimable ≡ old_free

    The per-collection table uses that expression for "live old gen the collection has to trace =
    182.2 MiB"
    , and its next row reports "old_free — free old-gen space at the trigger = 182.2
    MiB"
    . Those are the same quantity, and the source says which label is right —
    gc/old_free.rs:59: "Total bytes currently sitting in reusable old-gen holes." It is free
    space. The issue's own conclusion from that row — "It is not memory pressure: 182 MiB of old-gen is
    free every time"
    — is the correct reading of it.

    What the medians actually decode to

    quantity value source
    old_free (reusable holes) 182.2 MiB printed
    old_reclaimable = watched = live + unswept garbage 47.6 MiB printed
    ⇒ old_gen_in_use_bytes() = watched + free 229.8 MiB derived
    ⇒ old-gen live ≤ 47.6 MiB live ⊆ watched

    Consequences

    1. "traced : reclaimed = 23:1" is wrong — it divides free space by bytes reclaimed. The real
      bound is ≤ 47.6/7.9 ≈ 6:1, and lower to the extent the watched quantity is unswept garbage
      rather than live.
    2. "Quadratic in the live set" is the wrong cost model for this run. ~2.0 G instructions per
      collection over ≤47.6 MiB live is ≥42 instructions per live byte (~340 per 8-byte word) —
      far too high for marking. The profile says where it goes, and it is not marking:
      IncrementalSweepState::step 19.7 %, trace_heap_rewrite_slots 12.1 %,
      visit_gc_rewrite_slots 6.1 %, OldToYoungRememberedRebuildState::step 2.9 % — all
      proportional to old-gen SIZE (229.8 MiB), not to live.

    What is unaffected

    Everything measured directly: GC ≈ 88.4 %, 244 full collections, 242 from the absolute
    arm
    , old_threshold 48 MiB every time, baseline ~42 MiB, ~7.9 MiB reclaimed, and the re-arm
    mechanism — which I confirmed independently at the baseline setter (policy.rs:2179 sets
    baseline = old_in_use post-reclaim ≈ 42 MiB, so baseline < threshold is true again immediately).
    The fix direction is unchanged. Its justification is not.

    The question this opens, which I think is now the more important one

    Old-gen is 229.8 MiB holding ≤47.6 MiB live, and every collection re-pays a sweep and slot-rewrite
    over all of it.
    That makes pacing only one of two multipliers: how many collections and what
    each one costs
    . Fixing only the first leaves an O(heap) sweep running less often.

    For the composition measurement specifically

    The live old-gen set to account for is ≤47.6 MiB, not 182 MiB — which also bounds what can be
    shape records: 2.26 M records × ~32 B ≈ 72 MiB exceeds it, so either the shape slab is outside
    the traced heap or far fewer records are live than are minted.

    Confirmation pending

    One PERRY_GC_DIAG=1 run, checking old_reclaimable ≡ old_in_use − old_free holds per trigger. It
    must by construction; if it does not, my reading is wrong and this comment is void. I will post
    either way.

  4. proggeramlug commented on Sep 21, 2026

    @proggeramlug
    ContributorAuthor

    Confirmed empirically (242/242) — plus the arm split measured, and the shape slab answered

    Re-derived from this issue's own artefact (/root/matrix/tsc_gcdiag.err), so it is the same run.

    1. The correction holds — measured, not argued

    old_in_use − old_free == old_reclaimable in 242 of 242 triggers, zero violations, and the
    medians reproduce the table exactly:

    quantity median
    old_in_use 229.9 MiB
    old_free (reusable holes) 182.2 MiB
    old_reclaimable (watched) 47.6 MiB
    old_baseline 39.8 MiB
    old_band 32.0 MiB

    So old_in_use − old_reclaimable = 182.3 MiB is old_free. "23:1" and "182 MiB live traced"
    are withdrawn; old-gen live is ≤ 47.6 MiB.

    2. Which arm fires — per trigger, not inferred

    Evaluating the real predicate (input = old_reclaimable + external_side):

    count
    absolute arm only 231
    proportional arm only 11
    both true 0
    neither true 0
    retaining = true 0 / 242

    Zero both and zero neither — every trigger explained by exactly one arm, a check that could have
    failed. This issue's mechanism claim is confirmed: the absolute arm fires 95 % of the time and
    the RETAINING exemption never applies here.

    The ~4× inference is supported, first-order. Median watched growth since baseline is 8.8 MiB
    against a 32.0 MiB band, so removing the absolute arm paces at 32 MiB instead of 8.8 MiB:
    ≈3.6× fewer collections (242 → ~67), and ≥3.6× once the rising baseline widens the band.
    Still an estimate.

    3. Is the shape table's slab inside the traced heap? No — but it is a root set walked in full every full collection

    This decides whether #10868 touches GC, so, from source:

    • Outside the traced heap. shapes_store.rs stores records in
      Box<[UnsafeCell<ShapeRecord>; CHUNK_LEN]> chunks off a two-level page directory — Rust/mimalloc,
      not gc_malloc. So 2.26 M × 32 B ≈ 72 MiB is RSS, not traced live. That is consistent with
      the ≤47.6 MiB bound; had the slab been inside the heap, 72 MiB could not fit in 47.6 MiB.
    • But not invisible to GC. object::shapes::scan_shape_table_rekey_mut is a registered root
      scanner
      (gc/mod.rs:1027). Minors take a young-scoped fast path (fix(hir): recognize the Bun platform global in diagnostics #9754); a full collection
      takes the full walk
      , marking through and rewriting each record's keys word — "the word the
      collector marks through and rewrites in place"
      (shapes_store.rs:48).

    #10868 reaches GC through root scanning, not live-heap tracing: every full collection walks a
    root set proportional to shapes minted (~2.26 M), and there are 244 of them.
    The two defects
    multiply — by a different mechanism than "the live set is made of shape records".

    For the composition measurement: there is something to find, but it is the keys arrays
    (~1.1 M ArrayHeaders, in the heap, plausibly a large share of ≤47.6 MiB), not the records.

    4. What I think this reframes

    Pacing is one of two multipliers. The other is that a full collection sweeps and slot-rewrites
    229.9 MiB to find ≤47.6 MiB live — old-gen is 4.8× its live set — and additionally walks
    an O(#shapes) root set. Fixing only the trigger leaves an O(heap) sweep running less often. Why
    old-gen carries 182 MiB of holes, whether they are ever coalesced or returned, and why #9644's
    idle-time compaction (41.9 MB → 13.6 MB, free list to zero) is not holding here, is what I am
    reading next. Nothing proposed or built.

  5. proggeramlug commented on Sep 21, 2026

    @proggeramlug
    ContributorAuthor

    What each full collection actually traces: 82.6% keys arrays, ~15 MiB of program data

    PERRY_GC_CENSUS (the runtime's own classifier, at the mark-complete point of a synchronous full
    collection) on tscwork.ts plus gc() calls from a TypeScript before transformer, so one census
    lands mid-transpile with the AST live and one between transpiles. Output byte-identical to
    node --expose-gc (typescript 5.8.2 193520). Full detail on #10868 (comment 5767602463).

    mid-transpile, live heap 130.8 MiB objects bytes share
    array 703,549 109.92 MiB 84.1%
    ↳ shared keys arrays 682,146 107.99 MiB 82.6%
    object (AST nodes) 49,681 10.01 MiB 7.7%
    object_meta 41,353 5.36 MiB 4.1%
    string 50,112 2.91 MiB 2.2%
    closure 20,336 2.24 MiB 1.7%

    Between transpiles the live heap is 15.6 MiB and keys arrays are still 67.9% of it. The
    program's genuine live data is about 15 MiB; the rest of what the collector traces is
    object-model bookkeeping, nearly all of it one object kind.

    This makes the two levers multiply. Pacing (this issue) divides the number of collections —
    244 per transpile. Canonical identity (#10868) divides the size of each one, and 82.6% is the
    ceiling on that factor for this workload. Neither subsumes the other.

    On the root-scan channel that lane 14 identified — scan_shape_table_rekey_mut walking every
    shape record on each full collection to rewrite its keys word — I can price it from the same
    profile that produced the 88.4%: 0.95% of the run (scan_shape_table_rekey_mut 5,971 samples,
    ShapeTableInner::facts_remove 1,155, shape_keys_entry_is_minor_relevant 660, of 826,329). So
    the walk is real and happens 244 times, but it is a ~1% item, not a major term. The shape
    table's weight on this workload is overwhelmingly (a) RSS — shapes.descriptors 51.85 MiB +
    families 25.69 + by_facts 25.00 ≈ 102.5 MiB of Rust heap, more than the live GC heap — and
    (b) the keys arrays it keys itself by, which are in the traced heap and are the 82.6% above.

    shapes.families has 682,223 entries against 682,146 shared keys arrays: one family per array.
    The keys-array population is the distinct-shape population, which is why collapsing shape identity
    collapses the traced set.

    Two caveats. The census is taken at an explicit full collection mid-emit, not at one of the 244
    OldReclaim triggers, and it covers the whole heap while old_reclaimable is old-gen only — so
    the live totals here are not the trigger telemetry's ≤47.6 MiB. The ratio is the finding, and
    it holds at two very different moments (82.6% and 67.9%). And shape_keys_arrays counts only
    arrays flagged GC_FLAG_SHAPE_SHARED, so the true keys-array share is at least that.

  6. proggeramlug commented on Sep 22, 2026

    @proggeramlug
    ContributorAuthor

    Lever (iv) (PR #10931, 8d29d790b vs merge-base 0fa391529) moves this issue's numbers substantially, and the effect is mostly in GC, not in the mint path.

    PERRY_GC_CENSUS on a tsc transpile workload, both arms printing identical output (typescript 5.8.2 193520) at the same iteration count, four manual full collections each:

    Full-collection pause times, per collection:

    base head
    ms 49.0 / 14.8 / 49.7 / 15.9 7.5 / 4.4 / 8.6 / 4.3

    ~5.7× less time in a full collection, which follows directly from tracing 78.4% fewer live objects (870,628 → 188,250) and an 81.3% smaller traced heap (137.1 MB → 25.7 MB). RSS at peak-live fell 62.1% (610.2 MB → 231.4 MB).

    The reason this is worth noting on this issue specifically: an estimate that prices only the mint path puts the lever's value near the noise floor. That misses where it actually pays. 82.6% of the traced live heap was keys arrays (113.2 MB of 137.1 MB), so collapsing their identity removes most of what every full collection walks. The GC cost measured here was not a mint cost at all.

    Two cautions for anyone re-measuring:

    • maxrss will under-report this to the point of showing nothing. It is a high-water mark and cannot fall when live memory does. The first reading of this change said memory "barely moved" for exactly that reason. Read a census at a collection instead.
    • Side tables exceeded the traced heap on the base arm (161.6 MB vs 137.1 MB), so a heap-only accounting misses more than half the footprint. They fell to 20.8 MB (−87.1%).

    This does not by itself resolve the pacing behaviour described in the issue — the re-arming trigger is unchanged — but it moves the live set that the threshold is being compared against, which is the input the arm re-arms on. Worth re-taking the 244-collection count on the lever-(iv) tree before tuning the threshold, since the ~42 MiB live set this issue describes is a large part of what changed.

  7. proggeramlug commented on Oct 4, 2026

    @proggeramlug
    ContributorAuthor

    Root-caused and fixed by #11990: tsc's GC share came from a quadratic string build in TypeScript's createTextWriter (var output; … output += s), not from pacing. The untyped binding never got the in-place append, and every .length read marked the string shared, so each append copied the whole output (6.55 GB of 7.0 GB allocated at writeText). With #11990: tsc 186.94G → 28.69G instructions (−84.7%), task-clock 20.2 s → 3.05 s, peak RSS 358 → 268 MB, full collections 237 → 1. GC is now ~15% of tsc's profile with no dominant item. Remaining GC follow-ups: fold the remembered-set rebuild into the full's mark pass; #11549; a young large-object space.

  8. proggeramlug commented on Oct 5, 2026

    @proggeramlug
    ContributorAuthor

    Closing: fixed by #11990 (tsc −84.7% instructions, full collections 237 → 1; current main 28.7G instructions with 1 full collection, output matches node). Evidence above.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions