Skip to content

perf: a read site whose hot key is in the overflow region can NEVER latch megamorphic — it stays armed forever and keeps paying way compares that can never hit (#7753 measured that at +37%) #10863

Description

@proggeramlug

A property-read site whose hot key lives in the overflow (spill) region can never latch megamorphic, however many shapes it rotates. It stays armed for the life of the process and keeps running the polymorphic way compares on every read — which is exactly the cost #7753 built the latch to prevent.

#7753's own justification, in pic_prime_get:

Megamorphic. A rotation wider than the ways hold never hits one, so the compare sequence becomes pure cost — measured at +37% on a 7-shape site, against a 2.5x SPEEDUP on a 5-shape one. That asymmetry is the whole reason this state word exists: without it the ways pay well inside capacity and punish just past it, which is not a trade a compiler gets to make on the user's behalf.

If that +37% is right, it is being paid today on every such site.

Evidence

Two fixtures, identical except for which key is read — inline vs overflow. 64 shapes each, 30 M reads, PERRY_IC_DIAG=1, perry v0.5.1621+ on perrymaster.

Hot key INLINE (first slot) — the latch works:

primes=23345152 same_token=0 (0.0 %) new_token=23345152 (100.0 %) in_ways=0 (0.0 %)
  | way_state: fresh=11290 armed=214491 megamorphic=23119371

Hot key OVERFLOW (last slot) — the latch never engages:

primes=16886528 same_token=0 (0.0 %) new_token=16886528 (100.0 %) in_ways=0 (0.0 %)
  | way_state: fresh=2 armed=16886526 megamorphic=0
                                      ^^^^^^^^^^^^
own_overflow_primed=16886527 own_inline_primed=1

Sixteen million primes, every one a new_token with in_ways=0 — the ways never hit, which is the definition of the case the latch exists for — and megamorphic=0 throughout. The same rotation with an inline hot key latches on 99.1 % of primes.

Mechanism

pic_prime_get:

let prev_is_overflow = (prev_slot as u64) & u64::from(IC_SLOT_OVERFLOW_BIT) != 0;
let cascade = prev_tok != 0 && prev_tok != token && !prev_is_overflow;
...
if !cascade {
    return;          // <-- returns BEFORE the eviction-run counter
}
...
None => {            // no free way
    let run = (state >> 8) + 1;
    if run >= PIC_MEGAMORPHIC_EVICTIONS {
        ...
        c[PIC_WAY_STATE] = -PIC_LATCH_RETRY;   // the latch

The consecutive-eviction counter that drives the latch is only reached through the cascade. An overflow-encoded MRU slot suppresses the cascade, so the counter never advances and PIC_WAY_STATE never goes negative.

#9287 suppressed the cascade for encoded slots deliberately and documented why:

The emitted WAY path does not [test the overflow bit]: it computes obj + header + slot*8 directly, and an encoded slot there would be a wild load. Keep encoded slots out of the ways entirely; a polymorphic site rotating overflow shapes re-primes the MRU per shape, which is exactly the pre-#7753 behaviour.

Keeping such shapes out of the ways is correct and is not what this issue disputes. The unintended consequence is that they are also kept out of the latch — so the site keeps the armed state it earned during warm-up, and the emitted gate keeps reading and comparing ways that can never hit.

Why it matters beyond the microbenchmark

Overflow slots are not exotic: any object with more properties than the inline region holds has them, which is most configuration objects, AST nodes and ORM row objects. A site reading such a key across many shapes is the ordinary megamorphic case.

It is also a plausible component of the +77-per-megamorphic-miss that #10843 measured, and it is independent of the inherited-read work that surfaced it (#10834 / #10842 / #10860).

Repro

// 64 shapes, hot key in the OVERFLOW region: megamorphic=0 forever
const shapes = [];
for (let i = 0; i < 64; i++) { const t = {}; for (let j = 0; j <= i; j++) { t["f" + j] = j; } t.z = i; shapes.push(t); }
function run(n){ let h = 0; for (let i = 0; i < n; i++) { h = (h + shapes[i % 64].z) | 0; } return h; }

The control differs only in reading f0 (inline in every shape) instead of z:

for (let i = 0; i < 64; i++) { const t = { f0: i }; for (let j = 1; j <= i; j++) { t["g" + j] = j; } shapes.push(t); }

Fixtures on perrymaster at /root/m1/ (m64.ts, m64inline.ts).

Not taking this

Surfaced while testing a hook-skip predicate for the inherited-read cache, which I rejected for unrelated reasons. Filing so the measurement survives outside that PR thread.

Two directions for whoever picks it up, neither measured:

Activity

  1. proggeramlug commented on Sep 21, 2026

    @proggeramlug
    ContributorAuthor

    Fixed in #10896.

    pic_prime_get used one predicate for two questions and one return for both. Split them: evicted (a different shape displaced the MRU entry — the evidence the latch counts) is now separate from cascade (may the displaced entry be absorbed into a way — no, if its slot is overflow-encoded). A cascade suppressed for the overflow reason now advances the eviction run on an armed site and latches at the same threshold an inline rotation does. Nothing is written to a way on that path, so #9287's rule is untouched.

    Your reading of the code was right on every point I could check.

    Counts on m64.ts, predicted before the arm was built and committed to the branch first:

    before predicted measured
    armed 100.0 % < 2 %, est. 0.76 % 0.758 %
    megamorphic 0.0 % > 90 %, est. 97.7 % 96.97 %
    in_ways 0 0 0

    Cost, marginal instr/read (min of 3, fitted 500k→5M, flat): m64 757.98 → 737.69, −20.3 (−2.68 %). m64inline unchanged at 1352.85 → 1352.83; own3/w4/inh/inh3 bit-identical.

    Two things that did not come out the way the issue frames them, for the record:

    • The +37 % does not transfer. In these fixtures every read already misses into the full js_object_get_field_ic handler (~740 instructions), so the way block is a ~20-instruction slice of a large read — 2.7 %, not 37 %. perf(codegen,runtime): polymorphic property-read cache + arr.length short-circuit — interp.ts 3.96s → 2.39s #7753 measured it against a read the cache could otherwise serve.
    • Direction (b) would not have worked here. "Armed but every way empty" is not this site's state: the single shape whose key is at an inline slot (shapes[0]) does get cascaded into a way during warm-up, so the site is armed with one populated way that hits 1-in-64 while the other three compares can never hit. Only direction (a) reaches it.

    There is one honest regression: a site where no shape carries the key inline (so it never arms and the fix never fires) costs +4.8 instr/read (+0.64 %). A sabotage arm with the fix compiled out measures +4.2 on the same fixture, so it is codegen layout in this function, not the change — details and the six variants measured are in the PR.

  2. proggeramlug commented on Sep 21, 2026

    @proggeramlug
    ContributorAuthor

    Two follow-ups, both of which change how this issue should be sized.

    The +37 % in this issue's value estimate is wrong — correcting it here so nobody sizes work off it

    I opened this by quoting #7753's "+37 % on a 7-shape site" as the cost being paid today. It does not transfer to this bug and I could not reproduce anything like it. #7753 measured that against a read its cache could otherwise serve. At a site with an overflow hot key, every read already misses into the full js_object_get_field_ic handler, so the way-compare block is a ~20-instruction slice of a ~758-instruction read.

    Measured, marginal instructions per read (perf stat -e instructions:u, min of 3, fitted N=500k→5M, flat to 0.1):

    base after #10896
    m64.ts (64 shapes, hot key overflow) 757.98 737.69 −20.29, −2.68 %

    So the right figure for this issue is ≈2.7 %, not 37 % — a real fix worth making, and one that also removes a state the latch was built to prevent, but an order of magnitude smaller than the number I put at the top of the report. Anyone planning against "+37 % on every overflow site" should replan.

    What the read actually costs, for whoever picks up megamorphic reads

    Not scoping this, just leaving the number measured so it does not have to be re-derived. Same host, same method:

    program perry node
    m64.ts — megamorphic read, hot key in overflow 757.98 (737.69 after the fix) 119.64 (0.5–2M slope 115.24, 2–5M slope 121.84) 6.3× node
    ops4/own3.ts — monomorphic own-property read 123.00 8.07 15.2× node

    node is not flat on m64.js — those two slopes are why — so treat 119.64 as approximate. perry is flat to 0.1 instr/read on both.

    The latch fix removes 20 of those 758 instructions. The remaining ~740 is the prize and nobody owns it. It is six times what the same runtime spends on a monomorphic own-property read, and it is spent entirely inside the miss handler — shape lookup, key walk, overflow read, prime — on a site that is structurally going to miss, because the rotation is wider than any per-site cache is going to hold. Fixture is m64.ts in the repro above; PERRY_IC_DIAG attributes 100 % of it to own_overflow_primed.

    Worth noting the ratio inverts against the absolute: the megamorphic read is 6.3× node while the monomorphic one is 15.2×, so this is not the worst ratio in the runtime — it is the biggest absolute per-read figure I have measured on an ordinary property access.

  3. added 2 commits that reference this issue on Sep 21, 2026
  4. proggeramlug commented on Sep 21, 2026

    @proggeramlug
    ContributorAuthor

    Fixed in v0.5.1633 (merge train 253, 0fa3915293) by #10896.

    pic_prime_get used one predicate to answer two different questions — whether the displaced MRU entry may be absorbed into a polymorphic way, and whether a different shape displaced it at all — so a read site whose hot key lives in the overflow (spill) region could never reach the megamorphic latch.

    Verified by running crates/perry/tests/issue_9287_overflow_slot_ic.rs against the landed tree: 8 passed, 0 failed. The full perry-runtime suite is 4233/0/0 on the same tree.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions