Skip to content

perf: a dynamic string-keyed property read is 5.5x node when the key is present and 23x when absent, while static reads and array indexing both beat node #10753

Description

@proggeramlug

Summary

A dynamic string-keyed property read costs 5.5× node when the key is present and 23× when it is absent — while both of its component operations are things perry wins.

Per operation, fitted N=2,000→20,000, perf stat -e instructions:u. perry from origin/main + #10731 + #10746 + #10752; node v26.8.1; bun 1.3.14. Output checked equal to node on every row.

case perry node bun vs node
O.a — static key 113 153 56 1.35× — perry wins
W[i].length — the array index alone, no object read 70 202 105 2.88× — perry wins
O[K] — hoisted const K = "a" 622 190 50 0.30×
O[W[i]] — every key present 697 279 119 0.40×
O[W[i]] — some keys absent 3,118 135 150 0.04×
Map.get(W[i]) — same keys, for comparison 417 135 182 0.32×

What the decomposition says

Reading a statically known property is fast — perry beats node. Indexing the key array is fast — perry beats node by 2.9×. Composing them costs 697, against 183 for the two halves measured separately, so roughly 500 instructions appear that belong to neither.

And a miss is 4.5× worse than a hit — 3,118 against 697. That is the shape that matters most in practice, because O[k] || default, if (table[k]), and sparse lookup tables are all miss-heavy by design.

Note const K = "a" is barely better than a key read from an array (622 vs 697), so this is not about the key expression being dynamic — a constant string key already pays most of it. The cost is in the lookup itself.

Why this is worth doing next

It is larger than regex. On the per-op crossover table in #10695 regex is 0.15× against both runtimes; the miss path here is 0.04×, and unlike regex there is no narrower workaround — object property access by computed key is not an optional idiom.

Of the five realistic programs in #10695, records (0.45×) is object- and Set-bound, and tok (0.28×) does a keyword lookup per identifier. Both are in this path.

The related sibling, string-keyed Map, is #10697 — 4.2× node in the commonest shape, and Map.get measures 0.32× here, consistent with it.

Where I would look

The likely candidates, in the order I would check them:

  1. Per-lookup string hashing or interning — if the key's hash is recomputed per access rather than carried on the string, a constant key would pay it too, which matches const K costing 622.
  2. The miss path walking the prototype chain by name, with the 4.5× hit-to-miss ratio being that walk plus a second full lookup on Object.prototype. The JSON attribution in perf(json): JSON.stringify costs ~3,000 instructions per object visited (~20x node); array elements and primitives are fine #10696 found exactly this shape — a by-name prototype walk per object — costing ~450 instructions there.
  3. Whether an inline cache exists for computed keys at all. codegen: a string's .codePointAt in cc's hottest loop lowers to NativeMethodCall{module:"child_process", class_name:"Instance"} #9847 and the dynprop campaign made computed-key writes beat node (0.83×); reads appear not to have had the same treatment.

Related: #10695 (crossover map and real-program standings), #10697 (string-keyed Map), #10696 (the by-name prototype walk in JSON.stringify), #10741 (why primitive wins may not transfer to real loops — this one should, since it is a call-path cost rather than a loop-admission one).

Activity

  1. proggeramlug commented on Sep 19, 2026

    @proggeramlug
    ContributorAuthor

    Attribution: a miss walks the prototype chain by name, re-deriving the key at every level

    Callgrind (--dump-instr=yes, PERRY_TARGET_CPU=x86-64-v3, --debug-symbols), two programs identical except for which keys are present. Each performs 10,000 lookups; the miss program's keys are absent 6 times in 10.

    miss program : 54,971,882 Ir
    hit program  :  9,272,723 Ir
    

    A missing key costs ~7,600 instructions more than a hitting one. (45.70M extra ÷ 6,000 misses.)

    Where the miss program spends it

    Ir function
    3,063,440 object::native_get::try_data_get_bytes
    1,979,820 object::keys_lookup::keys_find_slot_by_bytes_resolved
    1,673,774 __memset_avx2_unaligned_erms (libc)
    1,667,997 js_object_get_field_by_name
    1,667,693 object::keys_lookup::keys_find_slot_by_bytes
    1,647,163 core::str::converts::from_utf8
    1,515,740 js_object_get_field_by_name'2
    1,406,490 object::class_registry::parent_static::is_class_object_ptr
    1,305,404 field_get_set::get_field_by_name_tail::get_field_by_name_object_tail'2
    1,284,005 field_get_set::get_field_by_name_tail::get_field_by_name_object_tail
    984,004 field_get_set::accessors::prototype_property_value_with_guard
    954,980 object::prototype_chain::meta_capable_object
    950,708 object::class_meta_registry::get_parent_class_id
    813,145 object::prototype_chain::object_static_prototype

    The '2 suffixes are the tell: js_object_get_field_by_name and get_field_by_name_object_tail each appear twice, once as themselves and once as a recursive instance. The lookup recurses up the prototype chain by name, and each level repeats the whole sequence.

    Three specific costs stand out:

    1. from_utf8 — UTF-8 validation of the key on every lookup, at every level. 1,647,163 Ir on the miss program against 200,147 on the hit program, an 8× difference that tracks the extra prototype levels a miss visits. The key is a string that already exists; validating its bytes per lookup is pure overhead.
    2. __memset_avx2 — a buffer zeroed per lookup, 1,397,512 Ir even on the all-hits program (~140/lookup). Something is being cleared on every property read.
    3. Two separate byte searches — keys_find_slot_by_bytes and keys_find_slot_by_bytes_resolved — together 3.6M Ir, plus try_data_get_bytes at 3.1M, the single largest entry.

    Why this is the same defect family as #10696

    The JSON.stringify attribution found a by-name prototype walk run per object costing ~450 instructions, where the answer was a property of the shape rather than the instance. This is the same shape one level down: a by-name walk per lookup, re-deriving per level what is a property of the object's class.

    Fix directions

    1. Do not re-validate the key. from_utf8 on a string that is already a valid JS string is wasted at every level; the byte view should be derived once per lookup at most, and ideally carried on the string.
    2. Negative caching / an inline cache for computed-key reads. codegen: a string's .codePointAt in cc's hottest loop lowers to NativeMethodCall{module:"child_process", class_name:"Instance"} #9847 and the dynprop campaign made computed-key writes beat node (0.83×); reads appear never to have had the same treatment, and the miss path in particular has no memo — every O[k] || default pays the full chain walk every time.
    3. Resolve class metadata once per lookup rather than per level. is_class_object_ptr, get_parent_class_id, meta_capable_object and object_static_prototype total ~4.1M Ir and are all functions of the object's class, not of the key.

    (1) looks smallest and is measurable on its own. I am taking this.

  2. proggeramlug commented on Sep 19, 2026

    @proggeramlug
    ContributorAuthor

    Narrowed to one function: a miss costs 10,617 instructions in get_field_by_name_object_tail

    Refining the attribution above with call counts. The fixture performs 10,000 lookups on {a:1,b:2,c:3} with keys ["a","zz","c","yy","b"] — so 6,000 hits and 4,000 misses.

    js_object_get_field_by_name
      → get_field_by_name_object_tail   ( 4,000x)  = 42,468,160 Ir   77.25% of the program
      → try_data_get_bytes              (10,000x)  =  4,416,000 Ir    8.03%
    

    The counts identify the paths exactly. try_data_get_bytes runs once per lookup at 442 instructions and resolves all 6,000 hits. get_field_by_name_object_tail runs exactly 4,000 times — once per miss — at 10,617 instructions each, and is 77% of the whole program.

    So the fast path is not the problem and the prototype walk is not evenly distributed: the entire cost is one function, reached only on a miss.

    What runs per miss

    Inside js_object_get_field_by_name, each of the 4,000 misses additionally pays a gauntlet of "is this a special receiver?" probes, all at 4,000 calls each:

    calls probe
    4,000 typedarray_props::typed_array_addr_from_value
    4,000 async_hooks::try_async_resource_property_dispatch
    4,000 date::is_date_cell_addr
    4,000 object::read_stub::read_stub_probe
    4,000 class_registry::parent_static::is_class_object_ptr
    4,000 core::f64::trunc
    8,000 class_meta_registry::get_parent_class_id (two per miss)
    8,000 core::str::converts::from_utf8 (two per miss)

    Each is cheap alone; the receiver is a plain object literal and every one of them answers "no". They are re-asked on every miss for the same receiver class.

    from_utf8 at two calls per miss confirms the key's bytes are re-validated per prototype level rather than once per lookup. Separately, try_data_get_bytes itself calls from_utf8 14,898 times for 10,000 lookups — about 1.5 per lookup.

    __memset_avx2_unaligned_erms is 1,673,774 Ir and is present on the all-hits program too (1,397,512), so something clears a buffer on every property read, hit or miss.

    This sharpens the fix directions

    The target is now a single function rather than "the lookup path":

    1. The special-receiver gauntlet is a function of the receiver's class, not of the key. Eight probes × 4,000 misses, all answering "no" for the same class every time. A per-class summary bit — the shape this codebase already uses for accessor keys via accessor_key_bits, and for toJSON in perf(json): JSON.stringify costs ~3,000 instructions per object visited (~20x node); array elements and primitives are fine #10696 — would collapse all of them.
    2. Negative caching. A miss on {a,b,c} for key "zz" is a property of (class, key) and cannot change without a shape or prototype mutation, both of which already bump epochs.
    3. Hoist from_utf8 out of the per-level path; two validations per miss of a string that is already a valid JS string.

    Given that try_data_get_bytes resolves hits in 442 instructions, a miss that resolved in a comparable budget would take this row from 0.04× node to roughly parity, and would also move #10697's Map-adjacent shapes and the records and tok programs in #10695.

  3. proggeramlug commented on Sep 19, 2026

    @proggeramlug
    ContributorAuthor

    A concrete design, using a mechanism this codebase already has

    There is no negative caching anywhere in the runtime (grep -riE "absent_cache|negative_cache|miss_cache|known_absent" returns nothing). But the exact pattern needed is already here and already consulted on this path.

    The existing precedent

    ObjectMeta::accessor_key_bits (object/mod.rs:1497) is a 64-bit Bloom summary keyed on key_bytes_hash(key) & 63. Setting is conservative (|= bit, descriptor_state.rs:825); a clear bit is a proof — it establishes there is no accessor for that key without any lookup. try_data_get_bytes already relies on it:

    // A clear Bloom bit proves no accessor for THIS key; collisions
    // conservatively decline.
    if meta.is_null() || (*meta).accessor_key_bits & accessor_bit == 0 {

    The proposal

    Add a sibling summary over the keys that are present on the object and everywhere on its prototype chain. A clear bit then proves the key is absent from the whole chain, and the read can return undefined immediately instead of entering get_field_by_name_object_tail.

    This is the correct polarity for the Bloom: a filter of present keys gives sound negative answers, and collisions merely fall through to today's walk — never a wrong result.

    • Consult it in js_object_get_field_by_name before dispatching to the tail, next to where accessor_key_bits is already read. One load and one test on a word the receiver's meta already provides.
    • Set a bit wherever a key becomes reachable: own-key addition, defineProperty, and prototype assignment (which must union the new prototype's summary, or clear to the conservative all-ones).
    • Clear to all-ones — meaning "prove nothing, take the walk" — on anything not cheaply summarised. Correctness then degrades to today's behaviour rather than to a wrong answer.

    Why the numbers say this is the right target

    try_data_get_bytes resolves a hit in 442 instructions. A miss costs 10,617 in the tail. A miss that resolved in a comparable budget takes this row from 0.04× node to roughly parity, and it is a single load and test on the common path.

    The eight per-class probes listed in the previous comment — typed_array_addr_from_value, try_async_resource_property_dispatch, is_date_cell_addr, read_stub_probe, is_class_object_ptr, get_parent_class_id, trunc — are all skipped as a consequence, because they sit behind the dispatch this avoids.

    The correctness obligations, stated up front

    The bit must be set before the key becomes observable, for: own-key add and delete (delete may leave the bit set — that is conservative and safe), defineProperty including accessors, Object.assign and spread, prototype assignment and Object.setPrototypeOf, Object.create, class-prototype method installation, and anything mutating a builtin prototype through globalThis. A proxy receiver must decline outright.

    Each of those needs a test that fails without the invalidation — the standard on this campaign, and the one that caught #10746's guard being unwitnessed while its GC stress passed 60 runs.

    VTABLE_GEN already exists and is bumped at the class-registration and parent-static sites, so a generation-keyed variant is available if a per-object summary proves too invasive.

  4. proggeramlug commented on Sep 19, 2026

    @proggeramlug
    ContributorAuthor

    Retraction: the node column in this issue is wrong, the per-miss figure is wrong, and the proposed design cannot work

    All four of my comments above need correcting. Taking them in order of how badly they mislead.

    1. The node column was measured while node was still warming up

    I fitted per-op costs as a two-size delta at N=2,000→20,000, believing that cancels startup. It cancels a constant. It does not cancel a warmup curve. The same fixture (O.a), same binaries, fitted at three ranges:

    fit range node perry
    2,000 → 20,000 140 113
    20,000 → 200,000 27 113
    200,000 → 2,000,000 10 113

    perry is flat to the instruction — it is AOT-compiled and has no warmup. node's cost falls 14× across the ranges because its JIT is still tiering up inside my measurement window.

    So the corrected table is:

    fixture perry node perry/node
    O.a static key 113.7 9.4 12.1× slower
    O[K] hoisted const key 623.1 12.1 51.5× slower
    O[W[i]] all present 697.7 143.0 4.9× slower
    O[W[i]] 40% absent 3120.3 167.3 18.7× slower
    Map.get 424.7 76.7 5.5× slower
    array-index control 70.5 21.2 3.3× slower

    The two rows I described as "perry wins" — the static read and the array-index control — are 12.1× and 3.3× slower than node. perry's own column is unchanged and matches to the decimal; only node's was wrong.

    2. The per-miss figure was contaminated by a one-time bootstrap, and divided by the wrong count

    bootstrap_prototype_addr is #[cold], caches on success, and appears in the profile with calls=1 costing 21,083,768 Ir — 38.35% of the miss program. It forces populate_global_this_builtins, the entire global-object bootstrap (#10686), charged to whichever lookup happens to run first. The hit program never triggers it; its entire 9.27M total is smaller than that one call.

    Two-size intercepts confirm it: 23.83M (miss) against 2.27M (hit), a difference of 21.57M.

    Corrected: a miss costs ~6,745 instructions, not 10,617. And my "~7,600 extra per miss" divided 45.70M by 6,000 when this issue's own second comment establishes 4,000 misses.

    3. There is no multi-level by-name walk inside the fast path

    My first comment said the lookup "recurses up the prototype chain by name, and each level repeats the whole sequence." try_data_get_bytes makes about one own-key lookup per invocation. The '2 recursion I pointed at is real but is not what I described.

    4. The proposed design cannot fire on its own benchmark

    I proposed a sibling Bloom field on ObjectMeta, beside accessor_key_bits. meta is NULL for a plain object literal — ObjectMeta's own documentation says "null for ordinary objects (the common case)" — and the existing read is guarded if meta.is_null() || ... precisely because of that. A field there cannot fire on {a:1,b:2,c:3}, which is this issue's own fixture. Forcing one into existence costs 120 bytes on a 40-byte object, GC-allocated, plus a traced edge each.

    The mechanism and the polarity were right; the storage was wrong.

    What actually survives

    • O[K] with a hoisted constant key costs 5.5× the identical read spelled O.a, and 51.5× node. That is a pure hit path with no invalidation obligations, and it is a larger and safer target than the miss path I filed this about.
    • The miss path is still slow (18.7× node), but a realistic fix takes it to roughly 4.5×, not parity — node's miss is ~196 instructions.
    • None of the five real programs in perf: crossover map — perry beats node/bun on 7 of 9 primitives at 100k iterations, and loses on regex and JSON #10695 performs a computed-key read on a plain object at all, let alone an absent one. Every property read in them is a static name or an array index, and every one hits. This fix would move one microbenchmark and nothing else.

    And the fix, if wanted, already exists

    promise/then_probe.rs (#7910) is a shipped, audited provable-negative for Get(obj, "then") on plain objects — the same problem, one key wide. It re-proves the receiver each probe, caches only the Object.prototype verdict under a signature of {proto_addr, keys_addr, keys_len, obj_flags, class_id, epoch, vtable_gen} computed by calling the real lookup once, fails closed, and ships PERRY_THENABLE_VERIFY=1 as its own sabotage harness.

    It also exposes the trap that would have sunk my design: field_set_by_name.rs:82 bumps prop_plan_epoch only for the literal key "toJSON". A negative memo keyed on that epoch alone would be silently wrong for Object.prototype.k = v; only the keys_addr/keys_len signature fields catch it.

    Generalising that, rather than designing new storage, is the right shape.

  5. proggeramlug commented on Sep 28, 2026

    @proggeramlug
    ContributorAuthor

    Progress: #11594 landed. The inherited-read cache now serves object-literal receivers, through the default Object.prototype link, and confirmed-absent keys, and it primes from by-name reads. Package workloads, instructions per iteration: moment/diff_duration −51%, lru-cache/churn −42%, lru-cache/ttl_mixed −35%, moment/parse_format −29%, dayjs −21…26%, validator −24%, jsonwebtoken/hs256 −22%, date-fns −18%, rate-limiter-flexible −11…13%. The trade is +0.4–3.1 MB peak RSS, approved by the owner.

    This is one slice of the prototype-chain read gap this issue tracks. Leaving it open for the remaining receiver kinds and the gap to Node.

  6. proggeramlug commented on Sep 30, 2026

    @proggeramlug
    ContributorAuthor

    Current target (2026-09-30): the computed-key read O[k] with the key present, now 5.6× Node

    Package impact (#11464)

    Bucket row 5, string ops / key transcoding on by-name lookups, is 5.1% of equal-weight excess. It is ≥5% in 12 packages. Share of each package's excess: cron 12, uuid 7, axios / commander / date-fns / dayjs / decimal.js / fastify / node-forge / qs / validator 6, and big.js / ioredis / lru-cache / moment / redis 5. That bucket also holds string concat and slice costs that are not this issue (uuid stringify.js:7, js_dynamic_string_or_number_add).

    The profile charges the key-conversion leaves themselves (from_utf8, copy_of_key, intern_dispatch_bytes, perry_string_ref_from_dispatch_id) to the calling lookup, usually bucket row 1 (property lookup). Their self cost, taken from each workload's top-20 self list (so a lower bound):

    workload key-conversion instr/iter % of excess
    cron/next_dates 48.7M 10.3
    date-fns/diff_interval 42k 7.3
    ioredis/set_get 98k 6.7
    rate-limiter-flexible/get_penalty 8.7k 6.3
    lru-cache/churn 13k 6.1
    moment/diff_duration 199k 5.8
    axios/post_json · qs/stringify_nested 1.56M · 803k 5.1 · 5.1
    dayjs/diff_startof · validator/batch 410k · 1.14M 4.5 · 4.3

    Named sites for this issue:

    • validator util/merge.js: for (key in defaults) … obj[key]. By-name reads are 20% and js_for_in_keys_stable_value is 11% (PROFILE.md).
    • cron/luxon DateTime.isValid: a 10.7M instr/iter chain whose dominant leaf is intern_dispatch_bytes.
    • lru-cache #set: intern_dispatch_bytes.

    What landed since the issue was filed

    Reproducer, fresh numbers (instructions per read)

    const O: Record<string, number> = { a: 1, b: 2, c: 3, d: 4, e: 5, f: 6, g: 7, h: 8 };
    const HIT = ["a","b","c","d","e","f","g","h"], MISS = ["a","x","c","y","e","z","g","w"];
    function hit(n: number)  { let s = 0; for (let i = 0; i < n; i++) s += O[HIT[i & 7]]; return s; }
    function miss(n: number) { let s = 0; for (let i = 0; i < n; i++) { const v = O[MISS[i & 7]]; s += v === undefined ? 1 : v; } return s; }
    // controls: s += O.c  and  const K = "c"; s += O[K]
    case Perry Node Perry/Node
    O.c (static) 27 22 1.2×
    O[K], const K = "c" 27 37 0.7×, now folded to the static read
    O[W[i & 7]], all present 948 171 5.6×
    O[W[i & 7]], half absent 653 193 3.4×

    The miss is no longer the problem: at 653 it is now cheaper than the hit, where it used to be 4.5× the hit. What remains is the present-key computed read, about 920 instructions more than the same key read statically.

    Acceptance target

    • Hit and miss rows ≤ 2× Node on this reproducer (≈ ≤ 340 / ≤ 390 instr per read at today's Node numbers).
    • No regression on the package benches for cron, validator, lru-cache, moment, dayjs, date-fns, ioredis, qs and axios.

    Package check (Linux, needs perf). Build with cargo build --release -p perry -p perry-runtime-static -p perry-stdlib-static, then run:

    (cd benchmarks/packages && npm ci --ignore-scripts)
    python3 scripts/package_bench.py compile --perry-bin-dir /tmp/pb --filter cron --filter validator --filter lru-cache --filter moment --filter dayjs --filter date-fns --filter ioredis --filter qs --filter axios
    python3 scripts/package_bench.py run --perry-bin-dir /tmp/pb --arms node,perry --modes instr --filter cron --filter validator --filter lru-cache --filter moment --filter dayjs --filter date-fns --filter ioredis --filter qs --filter axios --out /tmp/pb/instr.json

    Run it once on a base-commit build and once on the branch, and compare instructions per iteration. --filter is a workload-id substring, and control/* always runs. For attribution, run profile --callgraph on PERRY_KEEP_SYMBOLS=1 binaries (see benchmarks/packages/PROFILE.md).

    Fresh numbers were measured on origin/main 5fbc2c3 (v0.5.1654) with cargo build --release -p perry -p perry-runtime-static -p perry-stdlib-static. Binaries were compiled with PERRY_NO_AUTO_OPTIMIZE=1 and compared with Node 26.5.1 on Linux x86-64 using perf stat -e instructions:u. Each figure is the median of 3 runs at two sizes of N, with per-op = ΔI/ΔN. N1 is at least 200k operations (20k calls for the factory bench), which keeps most of Node's JIT warm-up inside the constant term. Output was byte-identical to Node on every row. The #11464 figures come from benchmarks/packages/profile/callgraph.{md,json}, measured at Perry 36420d2 with auto-optimize. node-forge was measured from 2febf42 binaries. "% of excess" means the share of a package's (Perry − Node) instructions per iteration.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions