Skip to content

perf: reads that resolve on the prototype chain (absent keys, inherited data) are 590–2,100× slower than Node (read IC caches own properties only) #10495

Description

@proggeramlug

Found by the package performance audit (real npm packages compiled from source, profiled against Node 26.5.1) and
re-measured on Perry 7661bc0 (v0.5.1589), Linux x64. A property read whose answer is not an own data property of
the receiver — an absent key (opts.ttl on {}), an inherited data property (this.DB on a prototype), or a
primitive's inherited property (s.constructor.name) — misses the read IC on every execution and re-walks the
prototype chain by name: 7,000–9,000 instructions per read, against ~40 for an own-property read.

Reproduction

bench.ts (23 lines):

// Reads that resolve on (or fall off) the prototype chain: absent keys, inherited data constants, String.prototype
const variant = process.argv[2] || "absent3"; const N = Number(process.argv[3] || "1000000");
class D { $L: string; $d: number; $y: number; constructor(v: number) { this.$L = "en"; this.$d = v; this.$y = v % 7; } clone() { return 1; } }
(D.prototype as any).DB = 26; (D.prototype as any).DM = 67108863; (D.prototype as any).DV = 67108864;
const objs: any[] = []; for (let i = 0; i < 64; i++) objs.push(new D(i));
const bags: any[] = [{ status: undefined }, { ttl: 5 }, {}, { start: 1 }];
const strs: any[] = ["foo@bar.com", "not-an-email", "https://example.com/p", "example"];
class Cache { ttl = 3; allowStale = false; updateAgeOnGet = false; v = 0;
  get(k: number, getOptions: any = {}) { const { allowStale = this.allowStale, updateAgeOnGet = this.updateAgeOnGet, status } = getOptions; this.v = (this.v + k + (allowStale ? 1 : 0) + (status ? 1 : 0)) | 0; return this.v; }
  plain(k: number) { this.v = (this.v + k + (this.allowStale ? 1 : 0)) | 0; return this.v; } }
const c = new Cache();
const V: Record<string, (n: number) => number> = {
  own3(n) { let a = 0; for (let i = 0; i < n; i++) { const o = objs[i & 63]; a += (o.$L ? 1 : 0) + (o.$d ? 1 : 0) + (o.$y ? 1 : 0); } return a; },
  absent3(n) { let a = 0; for (let i = 0; i < n; i++) { const o = objs[i & 63]; a += (o.$u ? 1 : 0) + (o.$offset ? 1 : 0) + (o.$x ? 1 : 0); } return a; },
  literal_absent3(n) { let a = 0; for (let i = 0; i < n; i++) { const o = bags[i & 3]; a += (o.ttl === undefined ? 1 : 0) + (o.start === undefined ? 1 : 0) + (o.size === undefined ? 1 : 0); } return a; },
  proto_data3(n) { let a = 0; for (let i = 0; i < n; i++) { const o = objs[i & 63]; a = ((a + o.DB) % o.DV) & o.DM; } return a; },
  str_ctor_name(n) { let a = 0; for (let i = 0; i < n; i++) { const s = strs[i & 3]; if (s.constructor.name !== "String") throw new TypeError("x"); a += s.length; } return a; },
  str_length(n) { let a = 0; for (let i = 0; i < n; i++) { const s = strs[i & 3]; a += s.length; } return a; },
  lru_get_opts(n) { let a = 0; for (let i = 0; i < n; i++) a = c.get(i); return a; },
  lru_plain(n) { let a = 0; for (let i = 0; i < n; i++) a = c.plain(i); return a; },
};
V[variant](N / 5 | 0); c.v = 0; const t0 = performance.now(); const cs = V[variant](N);
console.log(`variant=${variant} checksum=${cs} ms=${(performance.now() - t0).toFixed(2)}`);
PERRY_NO_AUTO_OPTIMIZE=1 perry compile bench.ts -o bench
for v in own3 absent3 literal_absent3 proto_data3 lru_plain lru_get_opts str_length str_ctor_name; do
  node bench.ts $v 1000000; ./bench $v 1000000; done

Measurements

Median of 3, shared host (loaded; instruction counts are the load-independent figure). N = 1,000,000; Perry
instructions per iteration = (whole-process instructions:u − 38 M startup) / 1.2 M (timed run + N/5 warm-up).
Node's loop is largely JIT-folded for these reads, so the Node-relative ratios are lower-bound-ish; the
control-vs-variant gap inside Perry is the mechanism's cost.

variant Node loop ms Perry loop ms ratio Perry instructions (per iter) Node wall ms Perry wall ms
own3 (control: 3 own reads) 6.3 26.2 4.2× 0.35 G (260) 143 64
absent3 (3 absent keys, class instance) 1.3 2,972.7 ~2,300× 31.74 G (26,400) 113 3,415
literal_absent3 (3 absent keys, object literals) 4.5 2,677.4 590× 26.69 G (22,200) 118 3,241
proto_data3 (3 inherited data reads o.DB/o.DV/o.DM) 1.4 1,084.4 790× 13.89 G (11,540) 83 1,423
lru_plain (control) 1.1 11.0 10× 0.24 G (170) 100 42
lru_get_opts (lru-cache get(k, opts = {}) destructuring with defaults) 1.1 2,283.7 ~2,100× 31.79 G (26,460) 82 2,792
str_length (control) 1.2 9.2 7.6× 0.17 G (110) 80 40
str_ctor_name (validator assertString) 23.2 867.4 37× 11.32 G (9,400) 130 1,031

Checksums identical. Per read: absent ≈ 7,400–8,800 instructions, inherited data ≈ 3,850, vs ≈ 40–80 for an own read.

Impact

Audit profiles, v0.5.1587 (shares of Perry CPU in the named workload):

  • lru-cache 11.5.2: get_field_ic_miss_impl 22.2 % of the whole run, 10.2 % of it prototype walks for absent keys —
    set(k, v, setOptions = {}) / get(k, getOptions = {}) destructure { ttl = this.ttl, start, status, … } from an
    options bag that usually lacks every key (objects-group report; combined with property adds ≈ 36 %).
  • decimal.js ≈ 17 %, jsonwebtoken ≈ 16 %, lodash ≈ 13 % (absent-key reads + adds, objects-group report).
  • dayjs 15.3 %, date-fns 12.2 % (differenceInHours 29.4 %), validator 16.0 % (isUUID 31.5 %), qs 5.5 %
    ("property GET IC miss", strings-group report). dayjs wrapper() reads instance.$u/instance.$offset (absent);
    validator assertString reads input.constructor.name on every entry point.
  • node-forge (jsbn): BigInteger.prototype.DB/DM/DV read as this.DB in the multiply/reduce loops, part of the
    9–11 % property bucket (numeric-group report).

Mechanism

  • get_field_ic_miss_impl (crates/perry-runtime/src/object/field_get_set/ic_miss.rs:540) primes the read PIC only
    when the key is found among the receiver's OWN keys: prime_get for an own inline slot (ic_miss.rs:980) or an own
    overflow slot (ic_miss.rs:924). Any other outcome falls through to an unprimed js_object_get_field_by_name(obj, key)
    (ic_miss.rs:995) (verified). There is no negative (absent) or holder (prototype) entry, so the next execution misses
    again.
  • That call reaches get_field_by_name_object_tail
    (crates/perry-runtime/src/object/field_get_set/get_field_by_name_tail.rs:7), which resolves the key by name through
    resolve_inherited_field / ordinary_object_prototype_property_value (get_field_by_name_tail.rs:1365-1393,
    accessors.rs:280) and prototype_property_value_with_guard (accessors.rs:218, a RuntimeHandleScope with three
    roots plus a recursive js_object_get_field_by_name on the prototype) (verified).
  • Profile of absent3 (perf record, symbols): get_field_ic_miss_impl 97 % inclusive → get_field_by_name_object_tail
    76 % → ordinary_object_prototype_property_value 44 %; self time spread over js_object_get_field_by_name 9 %,
    try_data_get_bytes 5 %, from_utf8 4 %, keys_find_slot_by_bytes 4 %, class-registry probes
    (class_decl_prototype_object, get_parent_class_id, lookup_prototype_method, is_class_object_ptr) and
    shape_is_url_search_params 7 % (verified by profile).
  • proto_data3 takes the same path (lookup_prototype_method, resolve_proto_chain_field_inner, SipHash of &str).
  • str_ctor_name goes get_field_ic_miss_impl → string_property_get_miss
    (crates/perry-runtime/src/string/char_ops.rs:164) → js_reflect_get on String.prototype, then .name on the
    String constructor through closure_get_dynamic_prop + get_accessor_descriptor + js_get_global_this_builtin_value
    (verified by profile: 70 % under string_property_get_miss, 44 % js_reflect_get).

What fast looks like

A prototype-chain-aware read IC: cache (receiver shape → holder object + slot) for inherited data and
(receiver shape → absent) for misses, guarded by a validity token for the shapes on the chain (invalidated when any
prototype on the chain gains/loses that key or is re-parented). Targets on this microbenchmark: absent3 and
literal_absent3 within 3× of own3 (≤ ~800 instructions/iteration, from 22,000–26,000); proto_data3 ≤ 500
instructions/iteration; lru_get_opts within 5× of lru_plain; str_ctor_name ≤ 1,000 instructions/iteration
(the String.prototype.constructor holder and the builtin name are both cacheable).

Notes

Activity

  1. added
    performanceRuntime, compile-time, build-size, or memory performance
    package-auditFound by the 2026 package audit: compiling real npm packages from source instead of native bindings
    on Sep 17, 2026
  2. proggeramlug commented on Sep 21, 2026

    @proggeramlug
    ContributorAuthor

    Re-measured on v0.5.1623 + #10846 (this issue's numbers are from v0.5.1589), and confirmed still open. I filed #10849 before finding this issue and have closed it as a duplicate; the parts that are new are below.

    perf stat -x, -e instructions:u, min of 3, fitted 200 k → 5 M, perry flat across both ranges, output identical to node on every row. node v26.8.1, bun /root/.bun/bin/bun, perrymaster x86-64.

    fixture receiver perry node bun perry / node
    hit_plain {x:0,a:1}, reads .a 116.0 11.9 10.0 10x
    miss_plain {x:0,a:1}, reads .zzz 7351.7 14.7 14.1 500x
    miss_class new K(), reads .zzz 8669.7 15.2 12.9 570x
    miss_create Object.create(P), reads .zzz 16388.4 15.4 14.9 1064x

    Three things this adds to the report above:

    1. The miss is 63x the HIT on the same object (116 vs 7352), so the whole of the gap is the runtime discovering the answer is undefined — not general read overhead.
    2. The receiver's chain shape more than doubles it. An Object.create receiver costs 16,388 against a plain object's 7,352. The chain is re-walked per read on the miss path exactly as it was on the inherited-HIT path before perf(runtime): give an INHERITED property read an inline-cache hit (1364 to 278 instructions) #10834.
    3. The verdict is now cheaply memoizable, which it was not when this was filed. "This (ShapeId, interned key) resolves to nothing" is a function of the receiver's shape and its chain, and perf(runtime): one validity word replaces the per-hop chain walk, and one flags word replaces two registry probes #10842 has just landed object::proto_validity — one global word that says when any object used as a prototype has changed structurally, checked with one load and one compare, covering a chain of any depth. A negative entry for a miss is the same object as the negative entry perf(runtime): give an INHERITED property read an inline-cache hit (1364 to 278 instructions) #10834 already records for a refused inherited read, and it can key on the same word.

    One warning worth inheriting from #10834 for whoever takes this: recording only the POSITIVE answers there made an accessor-on-prototype read 424 instructions per read slower than no cache at all, because the chain walk ran and was thrown away on every read. The refusals have to be recorded too, with the one exception of a refusal caused by a VALUE (an undefined or hole in a slot), which a plain store can change while transitioning no shape and bumping no epoch.

    Repro (.ts twin identical — no TypeScript syntax; R.x = i keeps the body loop-variant so node cannot delete it, and node's slope is 14.7, not 0):

    const N = Number(process.argv[2]);
    const R = {x:0, a:1};
    function run(n){ let h=0; for(let i=0;i<n;i++){ R.x=i; if(R.zzz===undefined) h++; } return h+R.x; }
    console.log(run(N));

    Fixtures and harness on perrymaster at /root/miss/ (miss_setup.sh, miss_measure.sh). Not taking this — flagging the re-measurement so it is not lost.

  3. proggeramlug commented on Sep 28, 2026

    @proggeramlug
    ContributorAuthor

    Progress: #11594 landed. The inherited-read cache now serves object-literal receivers, through the default Object.prototype link, and confirmed-absent keys, and it primes from by-name reads. Package workloads, instructions per iteration: moment/diff_duration −51%, lru-cache/churn −42%, lru-cache/ttl_mixed −35%, moment/parse_format −29%, dayjs −21…26%, validator −24%, jsonwebtoken/hs256 −22%, date-fns −18%, rate-limiter-flexible −11…13%. The trade is +0.4–3.1 MB peak RSS, approved by the owner.

    This is one slice of the prototype-chain read gap this issue tracks. Leaving it open for the remaining receiver kinds and the gap to Node.

  4. proggeramlug commented on Oct 3, 2026

    @proggeramlug
    ContributorAuthor

    Partial progress in #11796 (by-name reads from holder shapes): proto_data3 18.8k → 556, fn_missing_prop 20k → 167 instructions/op; Zod ×5000 −2.36%, qs −4.6%, tsc −0.34%. Leaving open for the rest.

  5. proggeramlug commented on Oct 4, 2026

    @proggeramlug
    ContributorAuthor

    More progress in #11885 (by-name stores, literal definitions from the shape, builtin globals as static sites, Function.prototype intrinsic entries): Zod −14.6%, qs −10.4%, commander −7.5%, tsc −0.01% (fulls 76=76), RSS lower on all four (n=5). Leaving open.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    package-auditFound by the 2026 package audit: compiling real npm packages from source instead of native bindingsperformanceRuntime, compile-time, build-size, or memory performance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions