Repository navigation
docs(codegen): record that the per-site concat cache is admitted only for counted-loop induction variables — it never fires on cc (not a defect) #9824
Description
Activity
Correction: nothing is broken. I had the premise wrong.
I filed this asking whether "the workload changed or something broke". Having
now read the admission gate and checked the benchmark, the answer is neither,
and the issue should be re-read with that correction in front.The gate is narrow by construction, and it is behaving correctly
concat_site_cache::try_lower_concat_site_cachedadmits a site only when- the left operand is a string literal, and
- the right operand is proven to lie in
0..=CONCAT_SITE_ADMIT_MAX(255)
— as a constant, as aLocalGetwhose loop-induction interval has
lo >= 0 && hi <= 255, or asx % Cwith a small constantC.
That is a loop-induction range-analysis requirement, not a threshold or an
outlining interaction. cc has noliteral + proven-small-induction-variable
concat sites, so the lowering correctly declines to fire.The
bench_object_propertyattribution stands — I checked itbenchmarks/suite/bench_object_property.tsis exactly the admitted shape, and
it is the gate's own doc-comment example:const FIELDS = 20; for (let j = 0; j < FIELDS; j++) { obj["field_" + j] = i * FIELDS + j; // literal + induction var, hi = 19 <= 255 }
So the mechanism fires there, the measured win is real, and no re-checking of
that attribution is needed. What does not follow is that it transfers. The
admitted shape — a string literal concatenated with a counted-loop induction
variable bounded under 256 — is characteristic of microbenchmarks and rare in
application code. cc performs ~8,600 concatenations per reply and not one of
them qualifies.One thing my
nmevidence proves more sharply than I saidnm -ureports static references, not executions. Because the fill arm's
js_string_concat_site_valuecall sits in the emitted instruction stream
whenever the diamond is lowered, its absence from the object proves the
lowering was never emitted at all for cc — not that it was emitted and
never taken. The runtime counter (site=0) is consistent with that but weaker
on its own. That distinction is the useful part of this issue.What is left worth doing
Not a fix — a fact that is currently unrecorded. Right now "does this
optimisation apply to real programs?" is answerable only by someone repeating
thisnmcheck. A per-mechanism record (the shapescripts/gc_rekeyed_key_tables.json
uses) that names each codegen fast path, states whether it is expected to fire
on the cc-parity bundle, and justifies the answer, would turn that from an
unknown into a documented one — and would catch the genuine regression case: a
gate that used to fire and stops.I would drop the "cache does not work" framing from the title if it stays open;
the accurate framing is "the per-site concat cache is admitted only for
counted-loop induction variables, so it does not apply to cc — recorded, not
broken."- changed the title
[-]perf(strings): the per-site concat cache (#9514) is entirely absent from the compiled claude-code binary — no call sites emitted[/-][+]docs(codegen): record that the per-site concat cache is admitted only for counted-loop induction variables — it never fires on cc (not a defect)[/+]on Sep 5, 2026
Found while counting executions at string-concat sites for the cc-performance
campaign. Filing rather than fixing, because it belongs to whoever owns #9514
and it bears on an attribution outside my lane.
Two independent measurements agree: the cache is not cold, it is not there
Runtime counter. I added a counter at the top of
js_string_concat_site_value, before the slot lookup — so it counts calls, nothits. One 400-character streamed reply through the offline mock-API rig, on the
compiled claude-code TUI:
js_string_concatandjs_string_concat_chainare called ~8,600 times perreply between them.
js_string_concat_site_valueis called zero times, inthree runs across two separately-compiled binaries.
Static check, which explains why. The symbol is not in the linked binary at
all:
Every other concat entry point is present; the site-cache one has been
dead-stripped, which it can only be if codegen emitted no call to it for
this module. So this is not a cache that misses — it is a cache with no call
sites in the workload.
(This is the same method that settled #9802:
nmon the compiled object saidthe outlined IC miss-handler was never referenced, and that turned out to be
the whole story there too.)
Why it is worth someone's time
CONCAT_SITE_SLOTStable, its GC lifecycle (gc/tests/concat_site.rs) andthe codegen pass are all carried for a path this workload never takes.
bench_object_propertybeatingnode was attributed to the per-site concat mechanism. That is a different
program and the mechanism may well fire there — but "it works" has now been
shown to be workload-dependent in a way nobody had measured, so the
benchmark deserves the same
nm/counter check before the attribution isrelied on again.
strength of a measured admission gate, so either cc's concat shapes never
matched the specialisation, or something in codegen stopped matching them.
Those want different fixes and the counter above distinguishes them cheaply.
What I did NOT determine
Whether cc's concat sites should match.
concat_site_cache.rsspecialisesprefix + <number>(its doc comment gives"field_" + jforj < 20), and Ihave not characterised what cc's ~8,600 concatenations per reply actually look
like — they may legitimately all be string+string, in which case the answer is
"expected, and the cache is simply for other programs" and this issue closes as
working-as-intended with a note. The counter to settle it is a split of
js_string_concat*calls by right-hand-side type.Reproducing
PERRY_ENUM_DIAG=<path>on the branch of #9823 reports the counters above.Binary: a full compile of
cli_2.1.112.jsfromperf/for-in-deferred-shadow-set(main
c7361c87c+ that diff), measured withsecret-tests/cc-permission-harness/stream_scale.pyat chunk 100, length 400.https://claude.ai/code/session_014UZWia6L37DpA93VLtNK9m