Repository navigation
perf: reads that resolve on the prototype chain (absent keys, inherited data) are 590–2,100× slower than Node (read IC caches own properties only) #10495
Description
Activity
- addedperformanceRuntime, compile-time, build-size, or memory performanceRuntime, compile-time, build-size, or memory performancepackage-auditFound by the 2026 package audit: compiling real npm packages from source instead of native bindingsFound by the 2026 package audit: compiling real npm packages from source instead of native bindings
on Sep 17, 2026 Re-measured on v0.5.1623 + #10846 (this issue's numbers are from v0.5.1589), and confirmed still open. I filed #10849 before finding this issue and have closed it as a duplicate; the parts that are new are below.
perf stat -x, -e instructions:u, min of 3, fitted 200 k → 5 M, perry flat across both ranges, output identical to node on every row. node v26.8.1, bun/root/.bun/bin/bun, perrymaster x86-64.fixture receiver perry node bun perry / node hit_plain{x:0,a:1}, reads.a116.0 11.9 10.0 10x miss_plain{x:0,a:1}, reads.zzz7351.7 14.7 14.1 500x miss_classnew K(), reads.zzz8669.7 15.2 12.9 570x miss_createObject.create(P), reads.zzz16388.4 15.4 14.9 1064x Three things this adds to the report above:
- The miss is 63x the HIT on the same object (116 vs 7352), so the whole of the gap is the runtime discovering the answer is
undefined— not general read overhead. - The receiver's chain shape more than doubles it. An
Object.createreceiver costs 16,388 against a plain object's 7,352. The chain is re-walked per read on the miss path exactly as it was on the inherited-HIT path before perf(runtime): give an INHERITED property read an inline-cache hit (1364 to 278 instructions) #10834. - The verdict is now cheaply memoizable, which it was not when this was filed. "This (ShapeId, interned key) resolves to nothing" is a function of the receiver's shape and its chain, and perf(runtime): one validity word replaces the per-hop chain walk, and one flags word replaces two registry probes #10842 has just landed
object::proto_validity— one global word that says when any object used as a prototype has changed structurally, checked with one load and one compare, covering a chain of any depth. A negative entry for a miss is the same object as the negative entry perf(runtime): give an INHERITED property read an inline-cache hit (1364 to 278 instructions) #10834 already records for a refused inherited read, and it can key on the same word.
One warning worth inheriting from #10834 for whoever takes this: recording only the POSITIVE answers there made an accessor-on-prototype read 424 instructions per read slower than no cache at all, because the chain walk ran and was thrown away on every read. The refusals have to be recorded too, with the one exception of a refusal caused by a VALUE (an
undefinedor hole in a slot), which a plain store can change while transitioning no shape and bumping no epoch.Repro (
.tstwin identical — no TypeScript syntax;R.x = ikeeps the body loop-variant so node cannot delete it, and node's slope is 14.7, not 0):const N = Number(process.argv[2]); const R = {x:0, a:1}; function run(n){ let h=0; for(let i=0;i<n;i++){ R.x=i; if(R.zzz===undefined) h++; } return h+R.x; } console.log(run(N));
Fixtures and harness on perrymaster at
/root/miss/(miss_setup.sh,miss_measure.sh). Not taking this — flagging the re-measurement so it is not lost.- The miss is 63x the HIT on the same object (116 vs 7352), so the whole of the gap is the runtime discovering the answer is
Progress: #11594 landed. The inherited-read cache now serves object-literal receivers, through the default
Object.prototypelink, and confirmed-absent keys, and it primes from by-name reads. Package workloads, instructions per iteration: moment/diff_duration −51%, lru-cache/churn −42%, lru-cache/ttl_mixed −35%, moment/parse_format −29%, dayjs −21…26%, validator −24%, jsonwebtoken/hs256 −22%, date-fns −18%, rate-limiter-flexible −11…13%. The trade is +0.4–3.1 MB peak RSS, approved by the owner.This is one slice of the prototype-chain read gap this issue tracks. Leaving it open for the remaining receiver kinds and the gap to Node.
- added 4 commits that reference this issue
on Oct 3, 2026 Partial progress in #11796 (by-name reads from holder shapes): proto_data3 18.8k → 556, fn_missing_prop 20k → 167 instructions/op; Zod ×5000 −2.36%, qs −4.6%, tsc −0.34%. Leaving open for the rest.
- added 5 commits that reference this issue
on Oct 4, 2026 More progress in #11885 (by-name stores, literal definitions from the shape, builtin globals as static sites, Function.prototype intrinsic entries): Zod −14.6%, qs −10.4%, commander −7.5%, tsc −0.01% (fulls 76=76), RSS lower on all four (n=5). Leaving open.
Found by the package performance audit (real npm packages compiled from source, profiled against Node 26.5.1) and
re-measured on Perry 7661bc0 (v0.5.1589), Linux x64. A property read whose answer is not an own data property of
the receiver — an absent key (
opts.ttlon{}), an inherited data property (this.DBon a prototype), or aprimitive's inherited property (
s.constructor.name) — misses the read IC on every execution and re-walks theprototype chain by name: 7,000–9,000 instructions per read, against ~40 for an own-property read.
Reproduction
bench.ts(23 lines):Measurements
Median of 3, shared host (loaded; instruction counts are the load-independent figure). N = 1,000,000; Perry
instructions per iteration = (whole-process
instructions:u− 38 M startup) / 1.2 M (timed run + N/5 warm-up).Node's loop is largely JIT-folded for these reads, so the Node-relative ratios are lower-bound-ish; the
control-vs-variant gap inside Perry is the mechanism's cost.
own3(control: 3 own reads)absent3(3 absent keys, class instance)literal_absent3(3 absent keys, object literals)proto_data3(3 inherited data readso.DB/o.DV/o.DM)lru_plain(control)lru_get_opts(lru-cacheget(k, opts = {})destructuring with defaults)str_length(control)str_ctor_name(validatorassertString)Checksums identical. Per read: absent ≈ 7,400–8,800 instructions, inherited data ≈ 3,850, vs ≈ 40–80 for an own read.
Impact
Audit profiles, v0.5.1587 (shares of Perry CPU in the named workload):
get_field_ic_miss_impl22.2 % of the whole run, 10.2 % of it prototype walks for absent keys —set(k, v, setOptions = {})/get(k, getOptions = {})destructure{ ttl = this.ttl, start, status, … }from anoptions bag that usually lacks every key (objects-group report; combined with property adds ≈ 36 %).
differenceInHours29.4 %), validator 16.0 % (isUUID31.5 %), qs 5.5 %("property GET IC miss", strings-group report). dayjs
wrapper()readsinstance.$u/instance.$offset(absent);validator
assertStringreadsinput.constructor.nameon every entry point.BigInteger.prototype.DB/DM/DVread asthis.DBin the multiply/reduce loops, part of the9–11 % property bucket (numeric-group report).
Mechanism
get_field_ic_miss_impl(crates/perry-runtime/src/object/field_get_set/ic_miss.rs:540) primes the read PIC onlywhen the key is found among the receiver's OWN keys:
prime_getfor an own inline slot (ic_miss.rs:980) or an ownoverflow slot (
ic_miss.rs:924). Any other outcome falls through to an unprimedjs_object_get_field_by_name(obj, key)(
ic_miss.rs:995) (verified). There is no negative (absent) or holder (prototype) entry, so the next execution missesagain.
get_field_by_name_object_tail(
crates/perry-runtime/src/object/field_get_set/get_field_by_name_tail.rs:7), which resolves the key by name throughresolve_inherited_field/ordinary_object_prototype_property_value(get_field_by_name_tail.rs:1365-1393,accessors.rs:280) andprototype_property_value_with_guard(accessors.rs:218, aRuntimeHandleScopewith threeroots plus a recursive
js_object_get_field_by_nameon the prototype) (verified).absent3(perf record, symbols):get_field_ic_miss_impl97 % inclusive →get_field_by_name_object_tail76 % →
ordinary_object_prototype_property_value44 %; self time spread overjs_object_get_field_by_name9 %,try_data_get_bytes5 %,from_utf84 %,keys_find_slot_by_bytes4 %, class-registry probes(
class_decl_prototype_object,get_parent_class_id,lookup_prototype_method,is_class_object_ptr) andshape_is_url_search_params7 % (verified by profile).proto_data3takes the same path (lookup_prototype_method,resolve_proto_chain_field_inner, SipHash of&str).str_ctor_namegoesget_field_ic_miss_impl→string_property_get_miss(
crates/perry-runtime/src/string/char_ops.rs:164) →js_reflect_getonString.prototype, then.nameon theStringconstructor throughclosure_get_dynamic_prop+get_accessor_descriptor+js_get_global_this_builtin_value(verified by profile: 70 % under
string_property_get_miss, 44 %js_reflect_get).What fast looks like
A prototype-chain-aware read IC: cache
(receiver shape → holder object + slot)for inherited data and(receiver shape → absent)for misses, guarded by a validity token for the shapes on the chain (invalidated when anyprototype on the chain gains/loses that key or is re-parented). Targets on this microbenchmark:
absent3andliteral_absent3within 3× ofown3(≤ ~800 instructions/iteration, from 22,000–26,000);proto_data3≤ 500instructions/iteration;
lru_get_optswithin 5× oflru_plain;str_ctor_name≤ 1,000 instructions/iteration(the
String.prototype.constructorholder and the builtinnameare both cacheable).Notes
own3260 instructions/iteration), so this is purely themiss path.
get_field_by_name_tailprobes 4 registries before reading the GcHeader it then switches on (1.8% ofpipeline_big) #7867 (probe family on the IC-miss tail), perf(runtime): gate the rare-subclass probes on the object property-miss path (asyncpipe 9.69x -> 1.76x node) #7795 (memoizedObject.prototypeon the miss path). The same missing holder/negative cache also shows up on function values —see perf: missing-property reads on functions are ~2,600× and
Object.hasOwn/getPrototypeOf40–90× slower than Node (Function.prototype re-resolved by name per call) #10497; the write-side counterpart (adds) is perf: adding a property witho.k = vis ~100× slower than Node (static-key write IC primes only overwrites; no add-transition cache) #10496.