Repository navigation
perf(gc): 88% of a real TypeScript transpile is GC — 244 full collections per transpileModule, because #7937's absolute old-reclaim arm re-arms after every collection (a ~42 MiB live set parked just under the 48 MiB threshold) #10928
Description
Activity
Correction: the live set is ~42 MiB, not ~182 MiB. The 182 MiB was the free list.
I mislabelled a field in the body above. The pacing diagnosis stands. The traced-to-reclaimed
ratio does not.From
gc/policy.rs:1832:fn old_gen_reclaimable_pressure_bytes() -> usize { crate::arena::old_gen_in_use_bytes().saturating_sub(super::old_free_bytes()) }
So in the
[gc-trigger]line,old_reclaimableis the occupied old gen (live plus garbage
since the last collection), andold_in_useis the arena's whole old-gen footprint (occupied plus
free list). I tookold_in_use − old_reclaimableto be live data. It is identicallyold_free: the
equalityold_in_use == old_free + old_reclaimableholds exactly in all 242 OldReclaim
triggers. I even printed both medians in the body as 182.2 MiB and did not notice they were the
same number.Corrected, median of 242 full collections:
was is live old gen each full collection traces (the post-reclaim old_baseline)182.2 MiB~42 MiB occupied old gen at the trigger ( old_reclaimable)47.6 MiB 47.6 MiB reclaimed per full collection 7.9 MiB 7.9 MiB traced : reclaimed 23 : 1~5 : 1 free-list space in the old-gen arena ( old_free)182.2 MiB 182.2 MiB — free, not live The correction tightens the pacing explanation rather than weakening it. The live set (~42 MiB)
sits just under the fixed 48 MiB threshold, so the headroom before the next "first crossing"
is only ~6–8 MiB. That is why a full collection fires on every ~7.9 MiB of growth: a live set parked
just under a fixed threshold is exactly the case a threshold-only trigger handles worst.Also struck from the body:
- "The ~182 MiB live set never enters either side of the comparison." Wrong. The live set is
the ~42 MiB baseline, and it enters asbaseline. The inference beside it still holds: with a
~42 MiB baseline the proportional band is its 32 MiB floor, so removing the absolute arm alone
gives roughly one full collection per ~32 MiB instead of per ~7.9 MiB, about 4× fewer.
Unaffected: 244 full collections per
transpileModule, 242 of 256 triggersOldReclaim, the arm
that fires, ~2.0 G instructions per full collection, GC ≈ 88.4% of the run, ~224–232× node.With ~42 MiB of live data at ~2.0 G instructions per full collection, each collection costs ~48
instructions per live byte. That is high for a mark over 42 MiB, and it matches the top symbol
beingIncrementalSweepState::step(19.7%, sweep, not mark): sweep cost scales with the
arena's extent — occupied plus the 182 MiB free list — not with the live set. That is a
second, separate lever, and I'm marking it as inference until measured.- "The ~182 MiB live set never enters either side of the comparison." Wrong. The live set is
- changed the title
[-]perf(gc): 88% of a real TypeScript transpile is GC — 244 full collections per transpileModule, because #7937's absolute old-reclaim arm re-arms after every collection (traces 182 MiB to reclaim 7.9 MiB, 23:1)[/-][+]perf(gc): 88% of a real TypeScript transpile is GC — 244 full collections per transpileModule, because #7937's absolute old-reclaim arm re-arms after every collection (a ~42 MiB live set parked just under the 48 MiB threshold)[/+]on Sep 21, 2026 Correction:
old_in_use − old_reclaimableisold_free— the "182 MiB live traced" and "23:1" rows are not supportedSource-derived, posted before my own diag run, because a composition measurement is being built
against the 182 MiB figure right now. Pending empirical confirmation (below); I will post the
confirmation or the retraction.The algebra
old_reclaimableis defined asold_in_use − old_free—gc/policy.rs:1832:/// Old-gen pressure the reclaim arms act on: block-offset in-use minus the /// swept holes the free list can already hand back (#7437). pub(super) fn old_gen_reclaimable_pressure_bytes() -> usize { crate::arena::old_gen_in_use_bytes().saturating_sub(super::old_free_bytes()) }
and
diag_sites.rs:46-57printsold_in_use,old_freeandold_reclaimablefrom exactly those
definitions. So, as an identity:old_in_use − old_reclaimable ≡ old_freeThe per-collection table uses that expression for "live old gen the collection has to trace =
182.2 MiB", and its next row reports "old_free— free old-gen space at the trigger = 182.2
MiB". Those are the same quantity, and the source says which label is right —
gc/old_free.rs:59: "Total bytes currently sitting in reusable old-gen holes." It is free
space. The issue's own conclusion from that row — "It is not memory pressure: 182 MiB of old-gen is
free every time" — is the correct reading of it.What the medians actually decode to
quantity value source old_free(reusable holes)182.2 MiB printed old_reclaimable= watched = live + unswept garbage47.6 MiB printed ⇒ old_gen_in_use_bytes()= watched + free229.8 MiB derived ⇒ old-gen live ≤ 47.6 MiB live ⊆ watched Consequences
- "traced : reclaimed = 23:1" is wrong — it divides free space by bytes reclaimed. The real
bound is ≤ 47.6/7.9 ≈ 6:1, and lower to the extent the watched quantity is unswept garbage
rather than live. - "Quadratic in the live set" is the wrong cost model for this run. ~2.0 G instructions per
collection over ≤47.6 MiB live is ≥42 instructions per live byte (~340 per 8-byte word) —
far too high for marking. The profile says where it goes, and it is not marking:
IncrementalSweepState::step19.7 %,trace_heap_rewrite_slots12.1 %,
visit_gc_rewrite_slots6.1 %,OldToYoungRememberedRebuildState::step2.9 % — all
proportional to old-gen SIZE (229.8 MiB), not to live.
What is unaffected
Everything measured directly: GC ≈ 88.4 %, 244 full collections, 242 from the absolute
arm,old_threshold48 MiB every time, baseline ~42 MiB, ~7.9 MiB reclaimed, and the re-arm
mechanism — which I confirmed independently at the baseline setter (policy.rs:2179sets
baseline = old_in_usepost-reclaim ≈ 42 MiB, sobaseline < thresholdis true again immediately).
The fix direction is unchanged. Its justification is not.The question this opens, which I think is now the more important one
Old-gen is 229.8 MiB holding ≤47.6 MiB live, and every collection re-pays a sweep and slot-rewrite
over all of it. That makes pacing only one of two multipliers: how many collections and what
each one costs. Fixing only the first leaves an O(heap) sweep running less often.For the composition measurement specifically
The live old-gen set to account for is ≤47.6 MiB, not 182 MiB — which also bounds what can be
shape records: 2.26 M records × ~32 B ≈ 72 MiB exceeds it, so either the shape slab is outside
the traced heap or far fewer records are live than are minted.Confirmation pending
One
PERRY_GC_DIAG=1run, checkingold_reclaimable ≡ old_in_use − old_freeholds per trigger. It
must by construction; if it does not, my reading is wrong and this comment is void. I will post
either way.- "traced : reclaimed = 23:1" is wrong — it divides free space by bytes reclaimed. The real
Confirmed empirically (242/242) — plus the arm split measured, and the shape slab answered
Re-derived from this issue's own artefact (
/root/matrix/tsc_gcdiag.err), so it is the same run.1. The correction holds — measured, not argued
old_in_use − old_free == old_reclaimablein 242 of 242 triggers, zero violations, and the
medians reproduce the table exactly:quantity median old_in_use229.9 MiB old_free(reusable holes)182.2 MiB old_reclaimable(watched)47.6 MiB old_baseline39.8 MiB old_band32.0 MiB So
old_in_use − old_reclaimable= 182.3 MiB isold_free. "23:1" and "182 MiB live traced"
are withdrawn; old-gen live is ≤ 47.6 MiB.2. Which arm fires — per trigger, not inferred
Evaluating the real predicate (input =
old_reclaimable + external_side):count absolute arm only 231 proportional arm only 11 both true 0 neither true 0 retaining = true0 / 242 Zero both and zero neither — every trigger explained by exactly one arm, a check that could have
failed. This issue's mechanism claim is confirmed: the absolute arm fires 95 % of the time and
the RETAINING exemption never applies here.The ~4× inference is supported, first-order. Median watched growth since baseline is 8.8 MiB
against a 32.0 MiB band, so removing the absolute arm paces at 32 MiB instead of 8.8 MiB:
≈3.6× fewer collections (242 → ~67), and ≥3.6× once the rising baseline widens the band.
Still an estimate.3. Is the shape table's slab inside the traced heap? No — but it is a root set walked in full every full collection
This decides whether #10868 touches GC, so, from source:
- Outside the traced heap.
shapes_store.rsstores records in
Box<[UnsafeCell<ShapeRecord>; CHUNK_LEN]>chunks off a two-level page directory — Rust/mimalloc,
notgc_malloc. So 2.26 M × 32 B ≈ 72 MiB is RSS, not traced live. That is consistent with
the ≤47.6 MiB bound; had the slab been inside the heap, 72 MiB could not fit in 47.6 MiB. - But not invisible to GC.
object::shapes::scan_shape_table_rekey_mutis a registered root
scanner (gc/mod.rs:1027). Minors take a young-scoped fast path (fix(hir): recognize the Bun platform global in diagnostics #9754); a full collection
takes the full walk, marking through and rewriting each record'skeysword — "the word the
collector marks through and rewrites in place" (shapes_store.rs:48).
#10868 reaches GC through root scanning, not live-heap tracing: every full collection walks a
root set proportional to shapes minted (~2.26 M), and there are 244 of them. The two defects
multiply — by a different mechanism than "the live set is made of shape records".For the composition measurement: there is something to find, but it is the keys arrays
(~1.1 MArrayHeaders, in the heap, plausibly a large share of ≤47.6 MiB), not the records.4. What I think this reframes
Pacing is one of two multipliers. The other is that a full collection sweeps and slot-rewrites
229.9 MiB to find ≤47.6 MiB live — old-gen is 4.8× its live set — and additionally walks
an O(#shapes) root set. Fixing only the trigger leaves an O(heap) sweep running less often. Why
old-gen carries 182 MiB of holes, whether they are ever coalesced or returned, and why #9644's
idle-time compaction (41.9 MB → 13.6 MB, free list to zero) is not holding here, is what I am
reading next. Nothing proposed or built.- Outside the traced heap.
What each full collection actually traces: 82.6% keys arrays, ~15 MiB of program data
PERRY_GC_CENSUS(the runtime's own classifier, at the mark-complete point of a synchronous full
collection) ontscwork.tsplusgc()calls from a TypeScriptbeforetransformer, so one census
lands mid-transpile with the AST live and one between transpiles. Output byte-identical to
node --expose-gc(typescript 5.8.2 193520). Full detail on #10868 (comment 5767602463).mid-transpile, live heap 130.8 MiB objects bytes share array 703,549 109.92 MiB 84.1% ↳ shared keys arrays 682,146 107.99 MiB 82.6% object (AST nodes) 49,681 10.01 MiB 7.7% object_meta 41,353 5.36 MiB 4.1% string 50,112 2.91 MiB 2.2% closure 20,336 2.24 MiB 1.7% Between transpiles the live heap is 15.6 MiB and keys arrays are still 67.9% of it. The
program's genuine live data is about 15 MiB; the rest of what the collector traces is
object-model bookkeeping, nearly all of it one object kind.This makes the two levers multiply. Pacing (this issue) divides the number of collections —
244 per transpile. Canonical identity (#10868) divides the size of each one, and 82.6% is the
ceiling on that factor for this workload. Neither subsumes the other.On the root-scan channel that lane 14 identified —
scan_shape_table_rekey_mutwalking every
shape record on each full collection to rewrite itskeysword — I can price it from the same
profile that produced the 88.4%: 0.95% of the run (scan_shape_table_rekey_mut5,971 samples,
ShapeTableInner::facts_remove1,155,shape_keys_entry_is_minor_relevant660, of 826,329). So
the walk is real and happens 244 times, but it is a ~1% item, not a major term. The shape
table's weight on this workload is overwhelmingly (a) RSS —shapes.descriptors51.85 MiB +
families25.69 +by_facts25.00 ≈ 102.5 MiB of Rust heap, more than the live GC heap — and
(b) the keys arrays it keys itself by, which are in the traced heap and are the 82.6% above.shapes.familieshas 682,223 entries against 682,146 shared keys arrays: one family per array.
The keys-array population is the distinct-shape population, which is why collapsing shape identity
collapses the traced set.Two caveats. The census is taken at an explicit full collection mid-emit, not at one of the 244
OldReclaimtriggers, and it covers the whole heap whileold_reclaimableis old-gen only — so
the live totals here are not the trigger telemetry's ≤47.6 MiB. The ratio is the finding, and
it holds at two very different moments (82.6% and 67.9%). Andshape_keys_arrayscounts only
arrays flaggedGC_FLAG_SHAPE_SHARED, so the true keys-array share is at least that.Lever (iv) (PR #10931,
8d29d790bvs merge-base0fa391529) moves this issue's numbers substantially, and the effect is mostly in GC, not in the mint path.PERRY_GC_CENSUSon a tsc transpile workload, both arms printing identical output (typescript 5.8.2 193520) at the same iteration count, four manual full collections each:Full-collection pause times, per collection:
base head ms 49.0 / 14.8 / 49.7 / 15.9 7.5 / 4.4 / 8.6 / 4.3 ~5.7× less time in a full collection, which follows directly from tracing 78.4% fewer live objects (870,628 → 188,250) and an 81.3% smaller traced heap (137.1 MB → 25.7 MB). RSS at peak-live fell 62.1% (610.2 MB → 231.4 MB).
The reason this is worth noting on this issue specifically: an estimate that prices only the mint path puts the lever's value near the noise floor. That misses where it actually pays. 82.6% of the traced live heap was keys arrays (113.2 MB of 137.1 MB), so collapsing their identity removes most of what every full collection walks. The GC cost measured here was not a mint cost at all.
Two cautions for anyone re-measuring:
maxrsswill under-report this to the point of showing nothing. It is a high-water mark and cannot fall when live memory does. The first reading of this change said memory "barely moved" for exactly that reason. Read a census at a collection instead.- Side tables exceeded the traced heap on the base arm (161.6 MB vs 137.1 MB), so a heap-only accounting misses more than half the footprint. They fell to 20.8 MB (−87.1%).
This does not by itself resolve the pacing behaviour described in the issue — the re-arming trigger is unchanged — but it moves the live set that the threshold is being compared against, which is the input the arm re-arms on. Worth re-taking the 244-collection count on the lever-(iv) tree before tuning the threshold, since the ~42 MiB live set this issue describes is a large part of what changed.
- added 8 commits that reference this issue
on Sep 24, 2026 Root-caused and fixed by #11990: tsc's GC share came from a quadratic string build in TypeScript's createTextWriter (
var output; … output += s), not from pacing. The untyped binding never got the in-place append, and every.lengthread marked the string shared, so each append copied the whole output (6.55 GB of 7.0 GB allocated at writeText). With #11990: tsc 186.94G → 28.69G instructions (−84.7%), task-clock 20.2 s → 3.05 s, peak RSS 358 → 268 MB, full collections 237 → 1. GC is now ~15% of tsc's profile with no dominant item. Remaining GC follow-ups: fold the remembered-set rebuild into the full's mark pass; #11549; a young large-object space.Closing: fixed by #11990 (tsc −84.7% instructions, full collections 237 → 1; current main 28.7G instructions with 1 full collection, output matches node). Evidence above.
Summary
On a real TypeScript workload, ~88% of perry's instructions are the garbage collector, and
the cause is a pacing arm: 244 full collections in one
ts.transpileModulecall, eachtracing ~182 MiB of live old-gen to reclaim ~7.9 MiB. The arm that fires is the #7937
absolute first-crossing arm of
old_reclaim_pressure_due. It is meant to fire once, when old-genfirst crosses 48 MiB, but it re-arms after every collection, because the post-reclaim
baseline (~42 MiB) never reaches the threshold. The result is the fixed-bytes cadence that
#7592's own comment says makes total major-GC work quadratic in the live set.
The pacing code is byte-identical between the measured build (v0.5.1631,
f88acdaa1) and currentmain (
0fa391529).The workload and the gap
/root/fr/work/tscwork.tson perrymaster (lane 14's): builds a 1,200-line TypeScript source andcalls
ts.transpileModuleon it 3 times (typescript 5.8.2). Complete runs, both outputtypescript 5.8.2 290280, identical to node.perf stat, stripped binary)perf recordperry / node ≈ 224–232×. The two perry counts come from independent binaries and methods and
agree within 3%. This is a different workload from the 9.7× reported in September, which was
tsc --noEmittype-checking (scanner/parser overlib.*.d.ts), timed by wall clock on macOSarm64.
Where the instructions go
perf record -e instructions:u -c 2000003on aPERRY_KEEP_SYMBOLS=1+PERRY_NO_AUTO_OPTIMIZE=1build, self time, bucketed by the demangled symbol's module path.Relative standard error per bucket ≈ 1/√samples.
perry_runtime::gc::*, arena, mimalloc)weakref::is_weak_target_trace_slot(GC tracing, placed under "other" by the first pass)perry_runtime::object::*, closures, proxy, symbols)quicksort<usize>0.62%, SipHash ~0.2% — no call graph was recorded)perry_fn_*,perry_method_*, closures)GC ≈ 88.4% of the run. perry's own compiled code — 5.1 G — is less than node's entire
run (7.13 G). A zero-cost object model would leave perry at ~215× node.
Top 15 symbols: 14 are GC and the 15th is the weakref trace slot.
IncrementalSweepState::step19.7%,
ValidPointerSet::contains11.6%,trace_heap_rewrite_slots8.1% + 4.0%,visit_gc_rewrite_slots6.1%,visit_gc_layout_slot_descriptors5.9%,remember_evacuated_old_to_young_slot4.8%,remembered_child_needs_tracking3.2%,ValidPointerSetBuilder::step3.2%,OldToYoungRememberedRebuildState::step2.9%,per_object_slot_mask2.6%, … Two notes so nobody chases the wrong thing:gc::verifyis themodule that holds the production forwarding rewrite and remembered-set rebuild, not a debug
verifier left on; and
ValidPointerSetis already a per-block object-start bitmap with a fencebinary search. It is hot because it is called for every traced pointer, not because it is slow.
Why: the collection count
PERRY_GC_DIAG=1(diagnostic only, no behaviour change) on one iteration:[gc-full]full collections[gc]collections, all kindssite=alloc_point kind=OldReclaimsafepoint kind=ArenaBytesbudgeted_start due[gc-copy-minor]Every one of the 242 OldReclaim triggers reports the same
old_threshold=50331648(48 MiB). At thetrigger:
old_reclaimable) at triggerold_baseline(post-reclaim value)old_in_use − old_reclaimable)old_free— free old-gen space at the triggerIt is not memory pressure: 182 MiB of old-gen is free every time.
The arm, from source (
gc/policy.rs:1881, unchanged on main)With the measured values —
old_in_use≈ 47.6–50 MiB ≥ 48 MiB,baseline≈ 42 MiB < 48 MiB, andRETAININGfalse because tsc's young generation is dying — the absolute arm is true everytime, so the proportional arm is never consulted. The #7937 comment on this very arm says "it is
the arm that actually fires", and exempts it only while the heap is RETAINING. After each full
collection the baseline resets to the post-reclaim value (~42 MiB). That is still below 48 MiB, so
the "first crossing" is re-armed, and the next ~7.9 MiB of watched growth crosses it again.
That is the shape #7592's comment on
gc_old_reclaim_growth_band_bytesdescribes and was writtento remove: "each full reclaim costs O(live), so a fixed-bytes cadence makes total major-GC work
quadratic in the live set." Here the fixed cadence arrives through the absolute arm instead of the
band.
One step further, and marked as inference rather than measurement. Removing the absolute arm
alone would not make the pacing proportional to this heap. The proportional arm's baseline is the
same ~42 MiB watched quantity, so its band is
max(32 MiB floor, 42/2)= the 32 MiB floor: a fullcollection per ~32 MiB of watched growth instead of per ~7.9 MiB, roughly 4× fewer. The ~182 MiB
live set never enters either side of the comparison. Whether the watched quantity
(
old_gen_reclaimable_pressure_bytes) is the right thing to pace on is the deeper question, and itis for the GC lane.
Relation to existing issues
minor collections (13 minors at ~0.8 G each,
run_copied_minor_attempt55% inclusive). Hereminors are 26 of 256 collections; the cost is full collections paced by OldReclaim. Same
family (O(live) work repeated too often), different trigger.
+diamond executions in this same run. Against 1,600 G thatis ~0.002–0.003% of the run. The attribution above is why: the program is almost entirely GC.
mint/transition, at 2.4% of the run.
Diagnosis only; nothing built or proposed as a patch. Artefacts on perrymaster:
/root/matrix/tsc_attr.data(perf record),/root/matrix/tsc_gcdiag.err(GC diag),/root/matrix/attribute.py.