Repository navigation
perf: a dynamic string-keyed property read is 5.5x node when the key is present and 23x when absent, while static reads and array indexing both beat node #10753
Description
Activity
Attribution: a miss walks the prototype chain by name, re-deriving the key at every level
Callgrind (
--dump-instr=yes,PERRY_TARGET_CPU=x86-64-v3,--debug-symbols), two programs identical except for which keys are present. Each performs 10,000 lookups; the miss program's keys are absent 6 times in 10.miss program : 54,971,882 Ir hit program : 9,272,723 IrA missing key costs ~7,600 instructions more than a hitting one. (45.70M extra ÷ 6,000 misses.)
Where the miss program spends it
Ir function 3,063,440 object::native_get::try_data_get_bytes1,979,820 object::keys_lookup::keys_find_slot_by_bytes_resolved1,673,774 __memset_avx2_unaligned_erms(libc)1,667,997 js_object_get_field_by_name1,667,693 object::keys_lookup::keys_find_slot_by_bytes1,647,163 core::str::converts::from_utf81,515,740 js_object_get_field_by_name'21,406,490 object::class_registry::parent_static::is_class_object_ptr1,305,404 field_get_set::get_field_by_name_tail::get_field_by_name_object_tail'21,284,005 field_get_set::get_field_by_name_tail::get_field_by_name_object_tail984,004 field_get_set::accessors::prototype_property_value_with_guard954,980 object::prototype_chain::meta_capable_object950,708 object::class_meta_registry::get_parent_class_id813,145 object::prototype_chain::object_static_prototypeThe
'2suffixes are the tell:js_object_get_field_by_nameandget_field_by_name_object_taileach appear twice, once as themselves and once as a recursive instance. The lookup recurses up the prototype chain by name, and each level repeats the whole sequence.Three specific costs stand out:
from_utf8— UTF-8 validation of the key on every lookup, at every level. 1,647,163 Ir on the miss program against 200,147 on the hit program, an 8× difference that tracks the extra prototype levels a miss visits. The key is a string that already exists; validating its bytes per lookup is pure overhead.__memset_avx2— a buffer zeroed per lookup, 1,397,512 Ir even on the all-hits program (~140/lookup). Something is being cleared on every property read.- Two separate byte searches —
keys_find_slot_by_bytesandkeys_find_slot_by_bytes_resolved— together 3.6M Ir, plustry_data_get_bytesat 3.1M, the single largest entry.
Why this is the same defect family as #10696
The
JSON.stringifyattribution found a by-name prototype walk run per object costing ~450 instructions, where the answer was a property of the shape rather than the instance. This is the same shape one level down: a by-name walk per lookup, re-deriving per level what is a property of the object's class.Fix directions
- Do not re-validate the key.
from_utf8on a string that is already a valid JS string is wasted at every level; the byte view should be derived once per lookup at most, and ideally carried on the string. - Negative caching / an inline cache for computed-key reads. codegen: a string's .codePointAt in cc's hottest loop lowers to NativeMethodCall{module:"child_process", class_name:"Instance"} #9847 and the dynprop campaign made computed-key writes beat node (0.83×); reads appear never to have had the same treatment, and the miss path in particular has no memo — every
O[k] || defaultpays the full chain walk every time. - Resolve class metadata once per lookup rather than per level.
is_class_object_ptr,get_parent_class_id,meta_capable_objectandobject_static_prototypetotal ~4.1M Ir and are all functions of the object's class, not of the key.
(1) looks smallest and is measurable on its own. I am taking this.
Narrowed to one function: a miss costs 10,617 instructions in
get_field_by_name_object_tailRefining the attribution above with call counts. The fixture performs 10,000 lookups on
{a:1,b:2,c:3}with keys["a","zz","c","yy","b"]— so 6,000 hits and 4,000 misses.js_object_get_field_by_name → get_field_by_name_object_tail ( 4,000x) = 42,468,160 Ir 77.25% of the program → try_data_get_bytes (10,000x) = 4,416,000 Ir 8.03%The counts identify the paths exactly.
try_data_get_bytesruns once per lookup at 442 instructions and resolves all 6,000 hits.get_field_by_name_object_tailruns exactly 4,000 times — once per miss — at 10,617 instructions each, and is 77% of the whole program.So the fast path is not the problem and the prototype walk is not evenly distributed: the entire cost is one function, reached only on a miss.
What runs per miss
Inside
js_object_get_field_by_name, each of the 4,000 misses additionally pays a gauntlet of "is this a special receiver?" probes, all at 4,000 calls each:calls probe 4,000 typedarray_props::typed_array_addr_from_value4,000 async_hooks::try_async_resource_property_dispatch4,000 date::is_date_cell_addr4,000 object::read_stub::read_stub_probe4,000 class_registry::parent_static::is_class_object_ptr4,000 core::f64::trunc8,000 class_meta_registry::get_parent_class_id(two per miss)8,000 core::str::converts::from_utf8(two per miss)Each is cheap alone; the receiver is a plain object literal and every one of them answers "no". They are re-asked on every miss for the same receiver class.
from_utf8at two calls per miss confirms the key's bytes are re-validated per prototype level rather than once per lookup. Separately,try_data_get_bytesitself callsfrom_utf814,898 times for 10,000 lookups — about 1.5 per lookup.__memset_avx2_unaligned_ermsis 1,673,774 Ir and is present on the all-hits program too (1,397,512), so something clears a buffer on every property read, hit or miss.This sharpens the fix directions
The target is now a single function rather than "the lookup path":
- The special-receiver gauntlet is a function of the receiver's class, not of the key. Eight probes × 4,000 misses, all answering "no" for the same class every time. A per-class summary bit — the shape this codebase already uses for accessor keys via
accessor_key_bits, and fortoJSONin perf(json): JSON.stringify costs ~3,000 instructions per object visited (~20x node); array elements and primitives are fine #10696 — would collapse all of them. - Negative caching. A miss on
{a,b,c}for key"zz"is a property of (class, key) and cannot change without a shape or prototype mutation, both of which already bump epochs. - Hoist
from_utf8out of the per-level path; two validations per miss of a string that is already a valid JS string.
Given that
try_data_get_bytesresolves hits in 442 instructions, a miss that resolved in a comparable budget would take this row from 0.04× node to roughly parity, and would also move #10697'sMap-adjacent shapes and therecordsandtokprograms in #10695.- The special-receiver gauntlet is a function of the receiver's class, not of the key. Eight probes × 4,000 misses, all answering "no" for the same class every time. A per-class summary bit — the shape this codebase already uses for accessor keys via
A concrete design, using a mechanism this codebase already has
There is no negative caching anywhere in the runtime (
grep -riE "absent_cache|negative_cache|miss_cache|known_absent"returns nothing). But the exact pattern needed is already here and already consulted on this path.The existing precedent
ObjectMeta::accessor_key_bits(object/mod.rs:1497) is a 64-bit Bloom summary keyed onkey_bytes_hash(key) & 63. Setting is conservative (|= bit,descriptor_state.rs:825); a clear bit is a proof — it establishes there is no accessor for that key without any lookup.try_data_get_bytesalready relies on it:// A clear Bloom bit proves no accessor for THIS key; collisions // conservatively decline. if meta.is_null() || (*meta).accessor_key_bits & accessor_bit == 0 {
The proposal
Add a sibling summary over the keys that are present on the object and everywhere on its prototype chain. A clear bit then proves the key is absent from the whole chain, and the read can return
undefinedimmediately instead of enteringget_field_by_name_object_tail.This is the correct polarity for the Bloom: a filter of present keys gives sound negative answers, and collisions merely fall through to today's walk — never a wrong result.
- Consult it in
js_object_get_field_by_namebefore dispatching to the tail, next to whereaccessor_key_bitsis already read. One load and one test on a word the receiver'smetaalready provides. - Set a bit wherever a key becomes reachable: own-key addition,
defineProperty, and prototype assignment (which must union the new prototype's summary, or clear to the conservative all-ones). - Clear to all-ones — meaning "prove nothing, take the walk" — on anything not cheaply summarised. Correctness then degrades to today's behaviour rather than to a wrong answer.
Why the numbers say this is the right target
try_data_get_bytesresolves a hit in 442 instructions. A miss costs 10,617 in the tail. A miss that resolved in a comparable budget takes this row from 0.04× node to roughly parity, and it is a single load and test on the common path.The eight per-class probes listed in the previous comment —
typed_array_addr_from_value,try_async_resource_property_dispatch,is_date_cell_addr,read_stub_probe,is_class_object_ptr,get_parent_class_id,trunc— are all skipped as a consequence, because they sit behind the dispatch this avoids.The correctness obligations, stated up front
The bit must be set before the key becomes observable, for: own-key add and
delete(delete may leave the bit set — that is conservative and safe),definePropertyincluding accessors,Object.assignand spread, prototype assignment andObject.setPrototypeOf,Object.create, class-prototype method installation, and anything mutating a builtin prototype throughglobalThis. Aproxyreceiver must decline outright.Each of those needs a test that fails without the invalidation — the standard on this campaign, and the one that caught #10746's guard being unwitnessed while its GC stress passed 60 runs.
VTABLE_GENalready exists and is bumped at the class-registration and parent-static sites, so a generation-keyed variant is available if a per-object summary proves too invasive.- Consult it in
Retraction: the node column in this issue is wrong, the per-miss figure is wrong, and the proposed design cannot work
All four of my comments above need correcting. Taking them in order of how badly they mislead.
1. The node column was measured while node was still warming up
I fitted per-op costs as a two-size delta at N=2,000→20,000, believing that cancels startup. It cancels a constant. It does not cancel a warmup curve. The same fixture (
O.a), same binaries, fitted at three ranges:fit range node perry 2,000 → 20,000 140 113 20,000 → 200,000 27 113 200,000 → 2,000,000 10 113 perry is flat to the instruction — it is AOT-compiled and has no warmup. node's cost falls 14× across the ranges because its JIT is still tiering up inside my measurement window.
So the corrected table is:
fixture perry node perry/node O.astatic key113.7 9.4 12.1× slower O[K]hoisted const key623.1 12.1 51.5× slower O[W[i]]all present697.7 143.0 4.9× slower O[W[i]]40% absent3120.3 167.3 18.7× slower Map.get424.7 76.7 5.5× slower array-index control 70.5 21.2 3.3× slower The two rows I described as "perry wins" — the static read and the array-index control — are 12.1× and 3.3× slower than node. perry's own column is unchanged and matches to the decimal; only node's was wrong.
2. The per-miss figure was contaminated by a one-time bootstrap, and divided by the wrong count
bootstrap_prototype_addris#[cold], caches on success, and appears in the profile withcalls=1costing 21,083,768 Ir — 38.35% of the miss program. It forcespopulate_global_this_builtins, the entire global-object bootstrap (#10686), charged to whichever lookup happens to run first. The hit program never triggers it; its entire 9.27M total is smaller than that one call.Two-size intercepts confirm it: 23.83M (miss) against 2.27M (hit), a difference of 21.57M.
Corrected: a miss costs ~6,745 instructions, not 10,617. And my "~7,600 extra per miss" divided 45.70M by 6,000 when this issue's own second comment establishes 4,000 misses.
3. There is no multi-level by-name walk inside the fast path
My first comment said the lookup "recurses up the prototype chain by name, and each level repeats the whole sequence."
try_data_get_bytesmakes about one own-key lookup per invocation. The'2recursion I pointed at is real but is not what I described.4. The proposed design cannot fire on its own benchmark
I proposed a sibling Bloom field on
ObjectMeta, besideaccessor_key_bits.metais NULL for a plain object literal —ObjectMeta's own documentation says "null for ordinary objects (the common case)" — and the existing read is guardedif meta.is_null() || ...precisely because of that. A field there cannot fire on{a:1,b:2,c:3}, which is this issue's own fixture. Forcing one into existence costs 120 bytes on a 40-byte object, GC-allocated, plus a traced edge each.The mechanism and the polarity were right; the storage was wrong.
What actually survives
O[K]with a hoisted constant key costs 5.5× the identical read spelledO.a, and 51.5× node. That is a pure hit path with no invalidation obligations, and it is a larger and safer target than the miss path I filed this about.- The miss path is still slow (18.7× node), but a realistic fix takes it to roughly 4.5×, not parity — node's miss is ~196 instructions.
- None of the five real programs in perf: crossover map — perry beats node/bun on 7 of 9 primitives at 100k iterations, and loses on regex and JSON #10695 performs a computed-key read on a plain object at all, let alone an absent one. Every property read in them is a static name or an array index, and every one hits. This fix would move one microbenchmark and nothing else.
And the fix, if wanted, already exists
promise/then_probe.rs(#7910) is a shipped, audited provable-negative forGet(obj, "then")on plain objects — the same problem, one key wide. It re-proves the receiver each probe, caches only theObject.prototypeverdict under a signature of{proto_addr, keys_addr, keys_len, obj_flags, class_id, epoch, vtable_gen}computed by calling the real lookup once, fails closed, and shipsPERRY_THENABLE_VERIFY=1as its own sabotage harness.It also exposes the trap that would have sunk my design:
field_set_by_name.rs:82bumpsprop_plan_epochonly for the literal key"toJSON". A negative memo keyed on that epoch alone would be silently wrong forObject.prototype.k = v; only thekeys_addr/keys_lensignature fields catch it.Generalising that, rather than designing new storage, is the right shape.
Progress: #11594 landed. The inherited-read cache now serves object-literal receivers, through the default
Object.prototypelink, and confirmed-absent keys, and it primes from by-name reads. Package workloads, instructions per iteration: moment/diff_duration −51%, lru-cache/churn −42%, lru-cache/ttl_mixed −35%, moment/parse_format −29%, dayjs −21…26%, validator −24%, jsonwebtoken/hs256 −22%, date-fns −18%, rate-limiter-flexible −11…13%. The trade is +0.4–3.1 MB peak RSS, approved by the owner.This is one slice of the prototype-chain read gap this issue tracks. Leaving it open for the remaining receiver kinds and the gap to Node.
Current target (2026-09-30): the computed-key read
O[k]with the key present, now 5.6× NodePackage impact (#11464)
Bucket row 5, string ops / key transcoding on by-name lookups, is 5.1% of equal-weight excess. It is ≥5% in 12 packages. Share of each package's excess: cron 12, uuid 7, axios / commander / date-fns / dayjs / decimal.js / fastify / node-forge / qs / validator 6, and big.js / ioredis / lru-cache / moment / redis 5. That bucket also holds string concat and slice costs that are not this issue (uuid
stringify.js:7,js_dynamic_string_or_number_add).The profile charges the key-conversion leaves themselves (
from_utf8,copy_of_key,intern_dispatch_bytes,perry_string_ref_from_dispatch_id) to the calling lookup, usually bucket row 1 (property lookup). Their self cost, taken from each workload's top-20 self list (so a lower bound):workload key-conversion instr/iter % of excess cron/next_dates 48.7M 10.3 date-fns/diff_interval 42k 7.3 ioredis/set_get 98k 6.7 rate-limiter-flexible/get_penalty 8.7k 6.3 lru-cache/churn 13k 6.1 moment/diff_duration 199k 5.8 axios/post_json · qs/stringify_nested 1.56M · 803k 5.1 · 5.1 dayjs/diff_startof · validator/batch 410k · 1.14M 4.5 · 4.3 Named sites for this issue:
- validator
util/merge.js:for (key in defaults) … obj[key]. By-name reads are 20% andjs_for_in_keys_stable_valueis 11% (PROFILE.md). - cron/luxon
DateTime.isValid: a 10.7M instr/iter chain whose dominant leaf isintern_dispatch_bytes. - lru-cache
#set:intern_dispatch_bytes.
What landed since the issue was filed
- perf(runtime): inherited-read cache serves object literals and absent keys (default Object.prototype link, confirmed ABSENT entries, by-name priming) #11594 (merged 09-28, after the bench(packages): Phase 3 attribution -- call chains, root-cause buckets, per-call floor #11464 profile): the inherited-read cache now serves object literals and confirmed-absent keys, and by-name reads prime it.
closure_get_dynamic_propuses canonical interned keys. Package deltas: moment −51/−29%, lru-cache −42/−35%, dayjs −26/−21%, validator −24%, jsonwebtoken/hs256 −22%, date-fns −18%, cron −5.4%. - The read-miss-front series perf(runtime): megamorphic reads confirm a slot guess by key atom; 'answerable by position' is a shape fact #11633, shapes: link-time static ShapeIds by content — runtime adoption + driver pass; REGISTERED_TYPED_SHAPES deleted (charter step 4) #11653, perf: one GC-leaf miss front per generic read site; POSBOUND (D3) #11657, perf: the read miss front takes the receiver as the fused test holds it (D4) #11658 and perf: inherited reads answered from the holder's shape (A1) #11681 (merged 09-29) also touches this path. I did not attribute the delta below to any single one of them.
Reproducer, fresh numbers (instructions per read)
const O: Record<string, number> = { a: 1, b: 2, c: 3, d: 4, e: 5, f: 6, g: 7, h: 8 }; const HIT = ["a","b","c","d","e","f","g","h"], MISS = ["a","x","c","y","e","z","g","w"]; function hit(n: number) { let s = 0; for (let i = 0; i < n; i++) s += O[HIT[i & 7]]; return s; } function miss(n: number) { let s = 0; for (let i = 0; i < n; i++) { const v = O[MISS[i & 7]]; s += v === undefined ? 1 : v; } return s; } // controls: s += O.c and const K = "c"; s += O[K]
case Perry Node Perry/Node O.c(static)27 22 1.2× O[K],const K = "c"27 37 0.7×, now folded to the static read O[W[i & 7]], all present948 171 5.6× O[W[i & 7]], half absent653 193 3.4× The miss is no longer the problem: at 653 it is now cheaper than the hit, where it used to be 4.5× the hit. What remains is the present-key computed read, about 920 instructions more than the same key read statically.
Acceptance target
- Hit and miss rows ≤ 2× Node on this reproducer (≈ ≤ 340 / ≤ 390 instr per read at today's Node numbers).
- No regression on the package benches for cron, validator, lru-cache, moment, dayjs, date-fns, ioredis, qs and axios.
Package check (Linux, needs
perf). Build withcargo build --release -p perry -p perry-runtime-static -p perry-stdlib-static, then run:(cd benchmarks/packages && npm ci --ignore-scripts) python3 scripts/package_bench.py compile --perry-bin-dir /tmp/pb --filter cron --filter validator --filter lru-cache --filter moment --filter dayjs --filter date-fns --filter ioredis --filter qs --filter axios python3 scripts/package_bench.py run --perry-bin-dir /tmp/pb --arms node,perry --modes instr --filter cron --filter validator --filter lru-cache --filter moment --filter dayjs --filter date-fns --filter ioredis --filter qs --filter axios --out /tmp/pb/instr.jsonRun it once on a base-commit build and once on the branch, and compare instructions per iteration.
--filteris a workload-id substring, andcontrol/*always runs. For attribution, runprofile --callgraphonPERRY_KEEP_SYMBOLS=1binaries (seebenchmarks/packages/PROFILE.md).Fresh numbers were measured on
origin/main5fbc2c3 (v0.5.1654) withcargo build --release -p perry -p perry-runtime-static -p perry-stdlib-static. Binaries were compiled withPERRY_NO_AUTO_OPTIMIZE=1and compared with Node 26.5.1 on Linux x86-64 usingperf stat -e instructions:u. Each figure is the median of 3 runs at two sizes of N, with per-op = ΔI/ΔN. N1 is at least 200k operations (20k calls for the factory bench), which keeps most of Node's JIT warm-up inside the constant term. Output was byte-identical to Node on every row. The #11464 figures come frombenchmarks/packages/profile/callgraph.{md,json}, measured at Perry 36420d2 with auto-optimize. node-forge was measured from 2febf42 binaries. "% of excess" means the share of a package's (Perry − Node) instructions per iteration.- validator
- added a commit that references this issue
on Oct 3, 2026
Summary
A dynamic string-keyed property read costs 5.5× node when the key is present and 23× when it is absent — while both of its component operations are things perry wins.
Per operation, fitted N=2,000→20,000,
perf stat -e instructions:u. perry fromorigin/main+ #10731 + #10746 + #10752; node v26.8.1; bun 1.3.14. Output checked equal to node on every row.O.a— static keyW[i].length— the array index alone, no object readO[K]— hoistedconst K = "a"O[W[i]]— every key presentO[W[i]]— some keys absentMap.get(W[i])— same keys, for comparisonWhat the decomposition says
Reading a statically known property is fast — perry beats node. Indexing the key array is fast — perry beats node by 2.9×. Composing them costs 697, against 183 for the two halves measured separately, so roughly 500 instructions appear that belong to neither.
And a miss is 4.5× worse than a hit — 3,118 against 697. That is the shape that matters most in practice, because
O[k] || default,if (table[k]), and sparse lookup tables are all miss-heavy by design.Note
const K = "a"is barely better than a key read from an array (622 vs 697), so this is not about the key expression being dynamic — a constant string key already pays most of it. The cost is in the lookup itself.Why this is worth doing next
It is larger than regex. On the per-op crossover table in #10695 regex is 0.15× against both runtimes; the miss path here is 0.04×, and unlike regex there is no narrower workaround — object property access by computed key is not an optional idiom.
Of the five realistic programs in #10695,
records(0.45×) is object- andSet-bound, andtok(0.28×) does a keyword lookup per identifier. Both are in this path.The related sibling, string-keyed
Map, is #10697 — 4.2× node in the commonest shape, andMap.getmeasures 0.32× here, consistent with it.Where I would look
The likely candidates, in the order I would check them:
const Kcosting 622.Object.prototype. The JSON attribution in perf(json): JSON.stringify costs ~3,000 instructions per object visited (~20x node); array elements and primitives are fine #10696 found exactly this shape — a by-name prototype walk per object — costing ~450 instructions there.Related: #10695 (crossover map and real-program standings), #10697 (string-keyed
Map), #10696 (the by-name prototype walk inJSON.stringify), #10741 (why primitive wins may not transfer to real loops — this one should, since it is a call-path cost rather than a loop-admission one).