Repository navigation
Conversation
The performance release, and it reaches docmd: the transform is about 14%
faster, with output byte for byte unchanged.
Measured on the transform itself, 2,000 paragraphs, JIT warmed, seven rounds,
both assemblies verified in bin before each run:
2.5.1 min 35,503 median 36,005 max 41,034
2.6.0 min 30,387 median 30,910 max 36,218
-14.4% -14.1% -11.7%
2.5.1 all [35503,35590,35664,36005,36215,36552,41034]
2.6.0 all [30387,30412,30518,30910,31209,31341,36218]
Every one of 2.6.0's bottom six rounds beats every one of 2.5.1's seven.
There is a mechanism, which is what was missing when a 37% "gain" on 2.5.1
turned out to be load. Six of the seven perf commits between the tags target
apply-templates, and bf25277 specifically speeds up a union of child kind
tests. docmd's paragraph templates select
"w:r | w:ins | w:hyperlink | w:sdt | w:fldSimple | w:smartTag" in four places,
once per paragraph, on documents of thousands of paragraphs.
END TO END THE GAIN IS NOT RESOLVABLE on this hardware, and that is a property
of the instrument rather than of the engine. CLI wall time also carries process
startup, OPC unzip, composite build, serialisation and file I/O, and this box
has a measured ~6% run-to-run spread. The corpus runs were 390,116 ms on 2.5.1
against 366,810 / 374,190 / 368,737 on 2.6.0: consistent in direction, too
close to the noise floor to quote.
TWO MEASUREMENTS I FIRST REPORTED WERE WRONG, both because an uncommitted pin
change followed `git checkout` onto main -- an uncommitted edit to a file that
is identical on both branches does not get reverted by switching. So a "2.5.1
re-baseline" and a "2.5.1" transform run were both 2.6.0. The benchmark printed
the assembly version it had loaded, which is the only reason it surfaced; the
numbers above come from runs that assert the expected version is in bin before
measuring. For an A/B the pin belongs in a commit, not the working tree.
Also verified
full suite 383 tests, 382 pass, 1 skip (unchanged)
build clean, 0 warnings, 0 errors
corpus output 13 of 13 byte-identical, across two runs
restore --locked-mode exit 0
transitive Core 2.0.0 -> 2.1.0, XQuery 2.5.1 -> 2.6.0
Byte-identical output alongside a speedup is the expected combination: every
perf commit changes how apply-templates walks children and what it allocates,
not what it emits.
The release also carries a security fix, bd57e34, so a cancelled transformation
stops inside XPath, regex matching and sorting. docmd passes
CancellationToken through ConvertAsync but has no caller that cancels today, so
this costs nothing now and is correct when batch mode arrives.
Core moving 2.0.0 -> 2.1.0 is transitive here and harmless. It is NOT harmless
for the phoenixml repo, which implements IContainer; whoever takes 2.6.0 there
should expect a breaking change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018wZEgtuzaZPswykiGEBxbz
Records the 14% transform gain with the measurement conditions, and why the end-to-end figure is not quotable on this hardware. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018wZEgtuzaZPswykiGEBxbz
|
On hold — waiting for the next engine release. 2.6.0 is verified good: ~14% faster on the transform, 13 of 13 corpus documents byte-identical, 383 tests passing. But the next version carries further performance and memory-footprint work, and verification costs the same whether it crosses one release or two — so this waits and goes in as one bump rather than two. Nothing here is wasted. When the next version lands, the comparison gets cheaper because the baselines now exist, measured on the same box with the same harness:
Two things to carry forward into that measurement:
Converted to draft rather than closed so the branch, the measurements and this reasoning stay in one place. |
|
Target: engine 2.7.0. Not yet published as of 2026-10-07. This draft holds the 2.6.0 bump. When 2.7.0 publishes, it supersedes this PR. 2.7.0 carries further performance and memory work, so one bump replaces two. Protocol for that test, fixed in advance:
|
|
Superseded by #49, which takes 2.7.0 instead. 2.6.0 leaves One thing measured here is worth carrying forward rather than discarding: 2.6.0 is the fastest The release this branch was held for never arrived: nothing between the 2.6.0 and 2.7.0 tags is |
The performance release, and it reaches docmd: the transform is about 14% faster, with output
byte for byte unchanged.
Measured on the transform itself
2,000 paragraphs, JIT warmed, seven rounds, each assembly verified in
binbefore its run:Every one of 2.6.0's bottom six rounds beats every one of 2.5.1's seven.
There is a mechanism, which is what was conspicuously missing when a 37% "gain" on 2.5.1
turned out to be machine load. Six of the seven
perf:commits between the tags targetapply-templates;bf25277speeds up a union of child kind tests specifically. docmd'sparagraph templates select
w:r | w:ins | w:hyperlink | w:sdt | w:fldSimple | w:smartTagin fourplaces, once per paragraph, on documents of thousands of paragraphs.
Why end-to-end can't show it
Not resolvable on this hardware, and that is the instrument's fault rather than the engine's.
CLI wall time also carries process startup, OPC unzip, composite build, serialisation and file
I/O, against a measured ~6% run-to-run spread on this box. Corpus runs: 390,116 ms on 2.5.1
against 366,810 / 374,190 / 368,737 on 2.6.0 — consistent in direction, too close to the noise
floor to quote.
Two of my own measurements were wrong first
Worth recording because the cause is a git property, not a typo. I edited the pin without
committing it, and an uncommitted change to a file identical on both branches is not reverted
by
git checkout— so the 2.6.0 pin followed me back tomain. A "2.5.1 re-baseline" and a"2.5.1" transform run were both actually 2.6.0, which first produced a false 5.5% win and then a
false "no difference".
It surfaced because the benchmark printed the assembly version it had loaded. The numbers above
come from runs that assert the expected version is in
binbefore measuring. For an A/B the pinbelongs in a commit, not the working tree.
Also verified
restore --locked-modeByte-identical output alongside a speedup is the expected combination: every perf commit changes
how
apply-templateswalks children and what it allocates, not what it emits.Notes
bd57e34) makes a cancelled transformation stop inside XPath, regex matchingand sorting. docmd threads a
CancellationTokenthroughConvertAsyncbut nothing cancelstoday — costs nothing now, correct when batch mode arrives.
phoenixmlrepo, which implementsIContainer— whoever takes 2.6.0 there should expect abreaking change.
🤖 Generated with Claude Code
https://claude.ai/code/session_018wZEgtuzaZPswykiGEBxbz