Repository navigation
perf: a read site whose hot key is in the overflow region can NEVER latch megamorphic — it stays armed forever and keeps paying way compares that can never hit (#7753 measured that at +37%) #10863
Description
Activity
Fixed in #10896.
pic_prime_getused one predicate for two questions and onereturnfor both. Split them:evicted(a different shape displaced the MRU entry — the evidence the latch counts) is now separate fromcascade(may the displaced entry be absorbed into a way — no, if its slot is overflow-encoded). A cascade suppressed for the overflow reason now advances the eviction run on an armed site and latches at the same threshold an inline rotation does. Nothing is written to a way on that path, so #9287's rule is untouched.Your reading of the code was right on every point I could check.
Counts on
m64.ts, predicted before the arm was built and committed to the branch first:before predicted measured armed100.0 % < 2 %, est. 0.76 % 0.758 % megamorphic0.0 % > 90 %, est. 97.7 % 96.97 % in_ways0 0 0 Cost, marginal instr/read (min of 3, fitted 500k→5M, flat):
m64757.98 → 737.69, −20.3 (−2.68 %).m64inlineunchanged at 1352.85 → 1352.83;own3/w4/inh/inh3bit-identical.Two things that did not come out the way the issue frames them, for the record:
- The +37 % does not transfer. In these fixtures every read already misses into the full
js_object_get_field_ichandler (~740 instructions), so the way block is a ~20-instruction slice of a large read — 2.7 %, not 37 %. perf(codegen,runtime): polymorphic property-read cache +arr.lengthshort-circuit — interp.ts 3.96s → 2.39s #7753 measured it against a read the cache could otherwise serve. - Direction (b) would not have worked here. "Armed but every way empty" is not this site's state: the single shape whose key is at an inline slot (shapes[0]) does get cascaded into a way during warm-up, so the site is armed with one populated way that hits 1-in-64 while the other three compares can never hit. Only direction (a) reaches it.
There is one honest regression: a site where no shape carries the key inline (so it never arms and the fix never fires) costs +4.8 instr/read (+0.64 %). A sabotage arm with the fix compiled out measures +4.2 on the same fixture, so it is codegen layout in this function, not the change — details and the six variants measured are in the PR.
- The +37 % does not transfer. In these fixtures every read already misses into the full
Two follow-ups, both of which change how this issue should be sized.
The +37 % in this issue's value estimate is wrong — correcting it here so nobody sizes work off it
I opened this by quoting #7753's "+37 % on a 7-shape site" as the cost being paid today. It does not transfer to this bug and I could not reproduce anything like it. #7753 measured that against a read its cache could otherwise serve. At a site with an overflow hot key, every read already misses into the full
js_object_get_field_ichandler, so the way-compare block is a ~20-instruction slice of a ~758-instruction read.Measured, marginal instructions per read (
perf stat -e instructions:u, min of 3, fitted N=500k→5M, flat to 0.1):base after #10896 m64.ts(64 shapes, hot key overflow)757.98 737.69 −20.29, −2.68 % So the right figure for this issue is ≈2.7 %, not 37 % — a real fix worth making, and one that also removes a state the latch was built to prevent, but an order of magnitude smaller than the number I put at the top of the report. Anyone planning against "+37 % on every overflow site" should replan.
What the read actually costs, for whoever picks up megamorphic reads
Not scoping this, just leaving the number measured so it does not have to be re-derived. Same host, same method:
program perry node m64.ts— megamorphic read, hot key in overflow757.98 (737.69 after the fix) 119.64 (0.5–2M slope 115.24, 2–5M slope 121.84) 6.3× node ops4/own3.ts— monomorphic own-property read123.00 8.07 15.2× node node is not flat on
m64.js— those two slopes are why — so treat 119.64 as approximate. perry is flat to 0.1 instr/read on both.The latch fix removes 20 of those 758 instructions. The remaining ~740 is the prize and nobody owns it. It is six times what the same runtime spends on a monomorphic own-property read, and it is spent entirely inside the miss handler — shape lookup, key walk, overflow read, prime — on a site that is structurally going to miss, because the rotation is wider than any per-site cache is going to hold. Fixture is
m64.tsin the repro above;PERRY_IC_DIAGattributes 100 % of it toown_overflow_primed.Worth noting the ratio inverts against the absolute: the megamorphic read is 6.3× node while the monomorphic one is 15.2×, so this is not the worst ratio in the runtime — it is the biggest absolute per-read figure I have measured on an ordinary property access.
Fixed in v0.5.1633 (merge train 253,
0fa3915293) by #10896.pic_prime_getused one predicate to answer two different questions — whether the displaced MRU entry may be absorbed into a polymorphic way, and whether a different shape displaced it at all — so a read site whose hot key lives in the overflow (spill) region could never reach the megamorphic latch.Verified by running
crates/perry/tests/issue_9287_overflow_slot_ic.rsagainst the landed tree: 8 passed, 0 failed. The fullperry-runtimesuite is 4233/0/0 on the same tree.
A property-read site whose hot key lives in the overflow (spill) region can never latch megamorphic, however many shapes it rotates. It stays
armedfor the life of the process and keeps running the polymorphic way compares on every read — which is exactly the cost #7753 built the latch to prevent.#7753's own justification, in
pic_prime_get:If that +37% is right, it is being paid today on every such site.
Evidence
Two fixtures, identical except for which key is read — inline vs overflow. 64 shapes each, 30 M reads,
PERRY_IC_DIAG=1, perry v0.5.1621+ on perrymaster.Hot key INLINE (first slot) — the latch works:
Hot key OVERFLOW (last slot) — the latch never engages:
Sixteen million primes, every one a
new_tokenwithin_ways=0— the ways never hit, which is the definition of the case the latch exists for — andmegamorphic=0throughout. The same rotation with an inline hot key latches on 99.1 % of primes.Mechanism
pic_prime_get:The consecutive-eviction counter that drives the latch is only reached through the cascade. An overflow-encoded MRU slot suppresses the cascade, so the counter never advances and
PIC_WAY_STATEnever goes negative.#9287 suppressed the cascade for encoded slots deliberately and documented why:
Keeping such shapes out of the ways is correct and is not what this issue disputes. The unintended consequence is that they are also kept out of the latch — so the site keeps the
armedstate it earned during warm-up, and the emitted gate keeps reading and comparing ways that can never hit.Why it matters beyond the microbenchmark
Overflow slots are not exotic: any object with more properties than the inline region holds has them, which is most configuration objects, AST nodes and ORM row objects. A site reading such a key across many shapes is the ordinary megamorphic case.
It is also a plausible component of the +77-per-megamorphic-miss that #10843 measured, and it is independent of the inherited-read work that surfaced it (#10834 / #10842 / #10860).
Repro
The control differs only in reading
f0(inline in every shape) instead ofz:Fixtures on perrymaster at
/root/m1/(m64.ts,m64inline.ts).Not taking this
Surfaced while testing a hook-skip predicate for the inherited-read cache, which I rejected for unrelated reasons. Filing so the measurement survives outside that PR thread.
Two directions for whoever picks it up, neither measured:
!cascadereturn, so an overflow rotation latches on the same evidence an inline one does (the ways stay empty either way, which is what Property access past the FIRST slot misses its cache: 3 ms vs 28 ms for the same object (bench_object_property, 2.6x) #9287 requires);