Repository navigation
Tombstone deletes lose a live object's keys array under evacuating GC — Object.keys() returns empty, fields read NaN (main, gc-stress red) #9200
Description
Activity
Update: a plausible fix tested and it does NOT fix this, plus a minimization that moves the suspect. Recording both so the next attempt doesn't repeat them.
The fix that failed
object/delete_rest.rs's clone+compact arm callsjs_array_allocand then copies from akeyspointer loaded before that allocation, republishing through an equally staleobj. Twenty lines above, the owned-tombstone fork does the same thing correctly, and its comment states the rule: "The allocation can collect. Reload the receiver, then recover its authoritative old keys edge instead of copying through the pre-collection raw addresses."I mirrored that — handle scope,
across_mutaround the alloc, reload both — built it, and the reproducer fails identically:after2: 103,NaN,309,NaN,keys3:empty, five runs byte-identical.That is still a real latent defect and I'd fix it on its own merits (a relocating collection during that alloc corrupts the clone, and reloading
keysalso repairsshape_index_migrate_after_delete, which keys onkeys as usizeafter the collector has already rekeyedindicesto the new address). But it is not this bug, and the reason is simple in hindsight: the fixture'sclass Row { a; b; c }haskey_count = 3 < 16, so the delete takes the owned-tombstone fork — the arm that was already correct — and never reaches the clone arm at all.Minimization moves the suspect to the dispatch tower
A reduced fixture — same class, same delete, same 200k-object churn,
Object.keysand field reads after each churn — passes underPERRY_GC_HEAP_LIMIT=8 PERRY_GC_FORCE_EVACUATE=1, on both perry and node:A keys1: a|c a= 2 c= 6 B keys1: a|c a= 2 c= 6 C keys1: a|c a= 2 c= 6So a tombstone delete followed by evacuating churn is not sufficient. What the original fixture adds is the #7142 machinery it was written for:
pickAlllives in another module,rows.map(r => r.pick())lowers to the class-id dispatch tower, and the tower routes to a proven-thisclone guarded by an inline pointer compare against@perry_class_keys_*.That reframes the symptom.
pick()isthis.a * 100 + this.c;NaNmeans the clone's fixed-slot loads ran on a receiver whose layout no longer matches — i.e. the inline keys-pointer compare wrongly succeeded for a receiver whose delete had replaced its keys array. Before the churn it correctly fails (theafter:line is right); only after an evacuating collection does it start passing when it should not.So the question is what happens to those two pointers across a relocating collection: the receiver's own (post-delete, owned) keys array, and the canonical
@perry_class_keys_*the guard compares against. One of them is being rewritten in a way that makes two arrays that were distinct before the collection compare equal after it — or the receiver's edge is being restored to the canonical array outright, which would also explainObject.keys()coming back empty rather than merely wrong.Where I'd point the next attempt
Instrument, don't reason — I've now had one confident code-reading hypothesis fail against this fixture. Print both pointers (receiver keys edge and the class-keys global) either side of the collection, in the failing configuration, and see which one moves where. The candidates that fit "distinct before, equal after" are the class-keys remembering path (
remember_class_keys_array/CLASS_KEYS_BY_ID), the shape descriptor'skeysrewrite in the metadata phase, andshape_keys_address_is_recycled's prune — which drops a descriptor whose keys address it judges dead, and would leave exactly an emptyObject.keys().Not a duplicate of #9234, and here is the control — since that issue implicates the same two knobs this reproducer uses (
PERRY_GC_HEAP_LIMIT=8 PERRY_GC_FORCE_EVACUATE=1) and describes a root scan running against an unbuilt stack-map index, which would free live objects silently. That is close enough to this symptom that it needed ruling out rather than assuming.Three runs of this reproducer, same binary and same configuration as the original report:
run 1 rc=0 stderr_bytes=0 run 2 rc=0 stderr_bytes=0 run 3 rc=0 stderr_bytes=0- No panic and no assert. Stack-map assert fires under PERRY_GC_HEAP_LIMIT=8 + FORCE_EVACUATE: a root-scan path still skips ensure_built() (#9191 gap) #9234's failure is gc: fail closed if the native root scan runs before the stack-map index is built #9182's fail-closed assert firing, with a message, after the program prints its correct result. This one exits
rc=0with zero bytes on stderr and the wrong answer already in stdout (after2: 103,NaN,309,NaN,keys3:empty). The two are distinguishable by inspection, not by argument. - The assert is present in this build (gc: fail closed if the native root scan runs before the stack-map index is built #9182 is on
mainat84185b5656), so the unguarded root-scan path Stack-map assert fires under PERRY_GC_HEAP_LIMIT=8 + FORCE_EVACUATE: a root-scan path still skips ensure_built() (#9191 gap) #9234 describes is not being taken here. If it were, this run would have aborted rather than answered wrongly. PERRY_OBJECT_TOMBSTONES=0clears this with the knobs held constant. A root-scan gap in the collector would not care about the tombstone switch.
So the attribution in the original report stands: this is the tombstone delete path, not the stack-map build.
Worth noting for anyone using those knobs on other work, though: #9234 means
PERRY_GC_HEAP_LIMIT=8 PERRY_GC_FORCE_EVACUATE=1can now produce a second, unrelated failure on a shutdown or final-collection path. When using that configuration to attribute anything, check stderr and the exit code, not just the output diff — a panic and a wrong answer are different findings and this configuration can now produce either.- No panic and no assert. Stack-map assert fires under PERRY_GC_HEAP_LIMIT=8 + FORCE_EVACUATE: a root-scan path still skips ensure_built() (#9191 gap) #9234's failure is gc: fail closed if the native root scan runs before the stack-map index is built #9182's fail-closed assert firing, with a message, after the program prints its correct result. This one exits
Bisect complete: this was mitigated, not fixed — and the mitigation names the mechanism.
The commit that made the reproducer pass is #9212 (
43e8b24e34), which containsfix(runtime): return tombstone deletes to opt-in (#9200). Current main'sobject_tombstone_deletes_enabled()reads:// #9200: keep tombstones opt-in until an evacuating-GC interaction // with class dispatch is fixed. The default-on route can restore a // deleted receiver to its canonical class shape after relocation, // making Object.keys() empty and fixed-slot reads return wrong data.
Full bisect chain over
84185b5656..89cde4ff14(26 commits), each step 5 runs underPERRY_GC_HEAP_LIMIT=8 PERRY_GC_FORCE_EVACUATE=1, byte-compared to node:commit result 84185b5656(original report)FAIL 10/10 — deterministic, re-verified as a control 0cc9e03532(#9204)FAIL 5/5 41eab746d3(#9205)PASS 43e8b24e34(#9212 — tombstones → opt-in)PASS 2f05b9de70(#9208),761dce1e4e(#9209), current mainPASS So the corruption is off the default path, not gone:
PERRY_OBJECT_TOMBSTONES=1on current main brings it back, and the 6.5× populated-delete win from #9038 is switched off in the default configuration until this is actually fixed.Taking the fix. The mitigation comment plus my earlier minimization narrow it well: the failure requires (1) a tombstone delete, (2) a relocating collection, and (3) the #7142 cross-module dispatch tower — the minimized fixture without the tower passes. "Restore a deleted receiver to its canonical class shape after relocation" matches that precisely: the suspect is the collection-time path that re-derives an object's shape/keys from its class id (which delete deliberately preserves) rather than from its published shape, so a post-delete receiver comes out of evacuation wearing the canonical pre-delete layout. The two earlier tested negatives stand: it is NOT the clone-arm rooting (fix built and disproven — the fixture takes the owned-tombstone fork,
key_count 3 < 16), and NOT #9234's unbuilt stack-map index (no assert fires; stderr is empty).- added a commit that references this issue
on Sep 4, 2026
Summary
On current
main(84185b5656), an object that has been through a tombstone delete can lose its entire keys array across an evacuating collection.Object.keys()returns the empty string and a previously-live field readsNaN.This is mine —
PERRY_OBJECT_TOMBSTONES=0makes it pass, so the tombstone delete path shipped in #9038 owns it. Filing rather than quietly fixing because it is onmain, it is a silent wrong-answer bug, andgc-stressis red for every PR because of it.Reproducer
Fixture already in the tree:
test-files/test_gap_repsel_pshape_tower_delete.ts.Deterministic — three runs byte-identical to each other, and byte-identical across an independent rebuild.
after2:103,206,309,NaN103,NaN,309,NaNkeys3:b|ckeys3isObject.keys(rows[1]).join("|")after a second delete plus allocation churn. An empty result means the receiver's keys array is gone, not that a key was removed — and the206 → NaNon the same object is the matching field read.Attribution matrix
Each row is the same binary, only the environment differs:
PERRY_GC_HEAP_LIMIT=8alonePERRY_GC_FORCE_EVACUATE=1alonePERRY_GC_HEAP_LIMIT=8+PERRY_GC_FORCE_EVACUATE=1PERRY_GC_HEAP_LIMIT=8+PERRY_GC_VERIFY_EVACUATION=1PERRY_GC_HEAP_LIMIT=8+FORCE_EVACUATE+VERIFY_EVACUATIONPERRY_GC_HEAP_LIMIT=8+FORCE_EVACUATE+VERIFY_EVACUATION+PERRY_OBJECT_TOMBSTONES=0Two things follow. The trigger needs both heap pressure and a relocating minor — neither alone reproduces, which is why the shipped default does not show it and why only the
force_verify/evac_minorfamily of arms is red. And the tombstone kill switch clears it, which localises the defect to the delete path rather than to the collector.PERRY_GC_VERIFY_EVACUATION=1does not fire, so this is not a slot the evacuation verifier covers — the stale reference is somewhere the verifier does not walk.Where I would look
The shape is "a raw reference held across a relocating collection" (#6993's family). The delete path clones the keys array during squeeze and republishes shape facts; the suspect set is any raw
keys/ObjectHeaderpointer live across an allocation inobject/delete_rest.rs'ssqueeze_holes_and_deleteand thepublish_object_shape_holescall ordering that #9110 already had to correct once (shape_drop → publish_object_shape_holes → set_object_live_slot_count).Note the symptom is the keys array being lost, not a wrong slot index — so a stale keys pointer surviving into the shape descriptor, or an object whose keys edge is not traced while it carries holes, fits better than an off-by-one in the squeeze.
Why this matters beyond the red gate
Silent, not a crash. An object quietly loses its properties, and the program continues with
undefined/NaN— the same failure class as #9192. It needs both heap pressure and relocation today, but the moving young-gen scavenge is the default collector since #7019, so the distance between this configuration and a real user's is smaller than the env vars make it look.Not fixed here
Recording the reproducer, the attribution and the deterministic diff. I am taking the fix.