Skip to content

Tombstone deletes lose a live object's keys array under evacuating GC — Object.keys() returns empty, fields read NaN (main, gc-stress red) #9200

Description

@proggeramlug

Summary

On current main (84185b5656), an object that has been through a tombstone delete can lose its entire keys array across an evacuating collection. Object.keys() returns the empty string and a previously-live field reads NaN.

This is mine — PERRY_OBJECT_TOMBSTONES=0 makes it pass, so the tombstone delete path shipped in #9038 owns it. Filing rather than quietly fixing because it is on main, it is a silent wrong-answer bug, and gc-stress is red for every PR because of it.

Reproducer

Fixture already in the tree: test-files/test_gap_repsel_pshape_tower_delete.ts.

perry compile test-files/test_gap_repsel_pshape_tower_delete.ts -o tower
PERRY_GC_HEAP_LIMIT=8 PERRY_GC_FORCE_EVACUATE=1 ./tower

Deterministic — three runs byte-identical to each other, and byte-identical across an independent rebuild.

line node 26.5.1 perry
after2: 103,206,309,NaN 103,NaN,309,NaN
keys3: b|c (empty)

keys3 is Object.keys(rows[1]).join("|") after a second delete plus allocation churn. An empty result means the receiver's keys array is gone, not that a key was removed — and the 206 → NaN on the same object is the matching field read.

Attribution matrix

Each row is the same binary, only the environment differs:

environment result
default (no env) PASS
PERRY_GC_HEAP_LIMIT=8 alone PASS
PERRY_GC_FORCE_EVACUATE=1 alone PASS
PERRY_GC_HEAP_LIMIT=8 + PERRY_GC_FORCE_EVACUATE=1 FAIL
PERRY_GC_HEAP_LIMIT=8 + PERRY_GC_VERIFY_EVACUATION=1 PASS
PERRY_GC_HEAP_LIMIT=8 + FORCE_EVACUATE + VERIFY_EVACUATION FAIL
PERRY_GC_HEAP_LIMIT=8 + FORCE_EVACUATE + VERIFY_EVACUATION + PERRY_OBJECT_TOMBSTONES=0 PASS

Two things follow. The trigger needs both heap pressure and a relocating minor — neither alone reproduces, which is why the shipped default does not show it and why only the force_verify/evac_minor family of arms is red. And the tombstone kill switch clears it, which localises the defect to the delete path rather than to the collector.

PERRY_GC_VERIFY_EVACUATION=1 does not fire, so this is not a slot the evacuation verifier covers — the stale reference is somewhere the verifier does not walk.

Where I would look

The shape is "a raw reference held across a relocating collection" (#6993's family). The delete path clones the keys array during squeeze and republishes shape facts; the suspect set is any raw keys / ObjectHeader pointer live across an allocation in object/delete_rest.rs's squeeze_holes_and_delete and the publish_object_shape_holes call ordering that #9110 already had to correct once (shape_drop → publish_object_shape_holes → set_object_live_slot_count).

Note the symptom is the keys array being lost, not a wrong slot index — so a stale keys pointer surviving into the shape descriptor, or an object whose keys edge is not traced while it carries holes, fits better than an off-by-one in the squeeze.

Why this matters beyond the red gate

Silent, not a crash. An object quietly loses its properties, and the program continues with undefined/NaN — the same failure class as #9192. It needs both heap pressure and relocation today, but the moving young-gen scavenge is the default collector since #7019, so the distance between this configuration and a real user's is smaller than the env vars make it look.

Not fixed here

Recording the reproducer, the attribution and the deterministic diff. I am taking the fix.

Activity

  1. proggeramlug commented on Aug 30, 2026

    @proggeramlug
    ContributorAuthor

    Update: a plausible fix tested and it does NOT fix this, plus a minimization that moves the suspect. Recording both so the next attempt doesn't repeat them.

    The fix that failed

    object/delete_rest.rs's clone+compact arm calls js_array_alloc and then copies from a keys pointer loaded before that allocation, republishing through an equally stale obj. Twenty lines above, the owned-tombstone fork does the same thing correctly, and its comment states the rule: "The allocation can collect. Reload the receiver, then recover its authoritative old keys edge instead of copying through the pre-collection raw addresses."

    I mirrored that — handle scope, across_mut around the alloc, reload both — built it, and the reproducer fails identically: after2: 103,NaN,309,NaN, keys3: empty, five runs byte-identical.

    That is still a real latent defect and I'd fix it on its own merits (a relocating collection during that alloc corrupts the clone, and reloading keys also repairs shape_index_migrate_after_delete, which keys on keys as usize after the collector has already rekeyed indices to the new address). But it is not this bug, and the reason is simple in hindsight: the fixture's class Row { a; b; c } has key_count = 3 < 16, so the delete takes the owned-tombstone fork — the arm that was already correct — and never reaches the clone arm at all.

    Minimization moves the suspect to the dispatch tower

    A reduced fixture — same class, same delete, same 200k-object churn, Object.keys and field reads after each churn — passes under PERRY_GC_HEAP_LIMIT=8 PERRY_GC_FORCE_EVACUATE=1, on both perry and node:

    A keys1: a|c a= 2 c= 6
    B keys1: a|c a= 2 c= 6
    C keys1: a|c a= 2 c= 6
    

    So a tombstone delete followed by evacuating churn is not sufficient. What the original fixture adds is the #7142 machinery it was written for: pickAll lives in another module, rows.map(r => r.pick()) lowers to the class-id dispatch tower, and the tower routes to a proven-this clone guarded by an inline pointer compare against @perry_class_keys_*.

    That reframes the symptom. pick() is this.a * 100 + this.c; NaN means the clone's fixed-slot loads ran on a receiver whose layout no longer matches — i.e. the inline keys-pointer compare wrongly succeeded for a receiver whose delete had replaced its keys array. Before the churn it correctly fails (the after: line is right); only after an evacuating collection does it start passing when it should not.

    So the question is what happens to those two pointers across a relocating collection: the receiver's own (post-delete, owned) keys array, and the canonical @perry_class_keys_* the guard compares against. One of them is being rewritten in a way that makes two arrays that were distinct before the collection compare equal after it — or the receiver's edge is being restored to the canonical array outright, which would also explain Object.keys() coming back empty rather than merely wrong.

    Where I'd point the next attempt

    Instrument, don't reason — I've now had one confident code-reading hypothesis fail against this fixture. Print both pointers (receiver keys edge and the class-keys global) either side of the collection, in the failing configuration, and see which one moves where. The candidates that fit "distinct before, equal after" are the class-keys remembering path (remember_class_keys_array / CLASS_KEYS_BY_ID), the shape descriptor's keys rewrite in the metadata phase, and shape_keys_address_is_recycled's prune — which drops a descriptor whose keys address it judges dead, and would leave exactly an empty Object.keys().

  2. proggeramlug commented on Aug 31, 2026

    @proggeramlug
    ContributorAuthor

    Not a duplicate of #9234, and here is the control — since that issue implicates the same two knobs this reproducer uses (PERRY_GC_HEAP_LIMIT=8 PERRY_GC_FORCE_EVACUATE=1) and describes a root scan running against an unbuilt stack-map index, which would free live objects silently. That is close enough to this symptom that it needed ruling out rather than assuming.

    Three runs of this reproducer, same binary and same configuration as the original report:

    run 1 rc=0 stderr_bytes=0
    run 2 rc=0 stderr_bytes=0
    run 3 rc=0 stderr_bytes=0
    

    So the attribution in the original report stands: this is the tombstone delete path, not the stack-map build.

    Worth noting for anyone using those knobs on other work, though: #9234 means PERRY_GC_HEAP_LIMIT=8 PERRY_GC_FORCE_EVACUATE=1 can now produce a second, unrelated failure on a shutdown or final-collection path. When using that configuration to attribute anything, check stderr and the exit code, not just the output diff — a panic and a wrong answer are different findings and this configuration can now produce either.

  3. proggeramlug commented on Aug 31, 2026

    @proggeramlug
    ContributorAuthor

    Bisect complete: this was mitigated, not fixed — and the mitigation names the mechanism.

    The commit that made the reproducer pass is #9212 (43e8b24e34), which contains fix(runtime): return tombstone deletes to opt-in (#9200). Current main's object_tombstone_deletes_enabled() reads:

    // #9200: keep tombstones opt-in until an evacuating-GC interaction
    // with class dispatch is fixed. The default-on route can restore a
    // deleted receiver to its canonical class shape after relocation,
    // making Object.keys() empty and fixed-slot reads return wrong data.

    Full bisect chain over 84185b5656..89cde4ff14 (26 commits), each step 5 runs under PERRY_GC_HEAP_LIMIT=8 PERRY_GC_FORCE_EVACUATE=1, byte-compared to node:

    commit result
    84185b5656 (original report) FAIL 10/10 — deterministic, re-verified as a control
    0cc9e03532 (#9204) FAIL 5/5
    41eab746d3 (#9205) PASS
    43e8b24e34 (#9212 — tombstones → opt-in) PASS
    2f05b9de70 (#9208), 761dce1e4e (#9209), current main PASS

    So the corruption is off the default path, not gone: PERRY_OBJECT_TOMBSTONES=1 on current main brings it back, and the 6.5× populated-delete win from #9038 is switched off in the default configuration until this is actually fixed.

    Taking the fix. The mitigation comment plus my earlier minimization narrow it well: the failure requires (1) a tombstone delete, (2) a relocating collection, and (3) the #7142 cross-module dispatch tower — the minimized fixture without the tower passes. "Restore a deleted receiver to its canonical class shape after relocation" matches that precisely: the suspect is the collection-time path that re-derives an object's shape/keys from its class id (which delete deliberately preserves) rather than from its published shape, so a post-delete receiver comes out of evacuation wearing the canonical pre-delete layout. The two earlier tested negatives stand: it is NOT the clone-arm rooting (fix built and disproven — the fixture takes the owned-tombstone fork, key_count 3 < 16), and NOT #9234's unbuilt stack-map index (no assert fires; stderr is empty).

  4. proggeramlug commented on Sep 2, 2026

    @proggeramlug
    ContributorAuthor

    Closing: the corruption was fixed structurally in #9317 (every post-birth ShapeId publish arms old_carrier) and tombstone deletes returned to default-on in #9331. Both are in release candidate r27.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions