Skip to content

deps: PhoenixmlDb.Xslt 2.5.1 -> 2.6.0 - #47

Closed
elvogel wants to merge 2 commits into
mainfrom
deps/engine-2.6.0
Closed

elvogel wants to merge 2 commits into
mainfrom
deps/engine-2.6.0

Conversation

@elvogel

@elvogel elvogel commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

The performance release, and it reaches docmd: the transform is about 14% faster, with output
byte for byte unchanged.

Measured on the transform itself

2,000 paragraphs, JIT warmed, seven rounds, each assembly verified in bin before its run:

2.5.1 2.6.0
min 35,503 ms 30,387 ms −14.4%
median 36,005 ms 30,910 ms −14.1%
max 41,034 ms 36,218 ms −11.7%
2.5.1  [35503, 35590, 35664, 36005, 36215, 36552, 41034]
2.6.0  [30387, 30412, 30518, 30910, 31209, 31341, 36218]

Every one of 2.6.0's bottom six rounds beats every one of 2.5.1's seven.

There is a mechanism, which is what was conspicuously missing when a 37% "gain" on 2.5.1
turned out to be machine load. Six of the seven perf: commits between the tags target
apply-templates; bf25277 speeds up a union of child kind tests specifically. docmd's
paragraph templates select w:r | w:ins | w:hyperlink | w:sdt | w:fldSimple | w:smartTag in four
places, once per paragraph, on documents of thousands of paragraphs.

Why end-to-end can't show it

Not resolvable on this hardware, and that is the instrument's fault rather than the engine's.
CLI wall time also carries process startup, OPC unzip, composite build, serialisation and file
I/O, against a measured ~6% run-to-run spread on this box. Corpus runs: 390,116 ms on 2.5.1
against 366,810 / 374,190 / 368,737 on 2.6.0 — consistent in direction, too close to the noise
floor to quote.

Two of my own measurements were wrong first

Worth recording because the cause is a git property, not a typo. I edited the pin without
committing it
, and an uncommitted change to a file identical on both branches is not reverted
by git checkout — so the 2.6.0 pin followed me back to main. A "2.5.1 re-baseline" and a
"2.5.1" transform run were both actually 2.6.0, which first produced a false 5.5% win and then a
false "no difference".

It surfaced because the benchmark printed the assembly version it had loaded. The numbers above
come from runs that assert the expected version is in bin before measuring. For an A/B the pin
belongs in a commit, not the working tree.

Also verified

Full suite 383 tests, 382 pass, 1 skip — unchanged
Build clean, 0 warnings / 0 errors
Corpus output 13 of 13 byte-identical, across two runs
restore --locked-mode exit 0
Transitive Core 2.0.0 → 2.1.0, XQuery 2.5.1 → 2.6.0

Byte-identical output alongside a speedup is the expected combination: every perf commit changes
how apply-templates walks children and what it allocates, not what it emits.

Notes

  • A security fix (bd57e34) makes a cancelled transformation stop inside XPath, regex matching
    and sorting. docmd threads a CancellationToken through ConvertAsync but nothing cancels
    today — costs nothing now, correct when batch mode arrives.
  • Core 2.0.0 → 2.1.0 is transitive and harmless here. It is not harmless for the
    phoenixml repo, which implements IContainer — whoever takes 2.6.0 there should expect a
    breaking change.

🤖 Generated with Claude Code

https://claude.ai/code/session_018wZEgtuzaZPswykiGEBxbz

Lucas Vogel and others added 2 commits October 6, 2026 10:08
The performance release, and it reaches docmd: the transform is about 14%
faster, with output byte for byte unchanged.

Measured on the transform itself, 2,000 paragraphs, JIT warmed, seven rounds,
both assemblies verified in bin before each run:

  2.5.1   min 35,503   median 36,005   max 41,034
  2.6.0   min 30,387   median 30,910   max 36,218
          -14.4%          -14.1%         -11.7%

  2.5.1 all [35503,35590,35664,36005,36215,36552,41034]
  2.6.0 all [30387,30412,30518,30910,31209,31341,36218]

Every one of 2.6.0's bottom six rounds beats every one of 2.5.1's seven.

There is a mechanism, which is what was missing when a 37% "gain" on 2.5.1
turned out to be load. Six of the seven perf commits between the tags target
apply-templates, and bf25277 specifically speeds up a union of child kind
tests. docmd's paragraph templates select
"w:r | w:ins | w:hyperlink | w:sdt | w:fldSimple | w:smartTag" in four places,
once per paragraph, on documents of thousands of paragraphs.

END TO END THE GAIN IS NOT RESOLVABLE on this hardware, and that is a property
of the instrument rather than of the engine. CLI wall time also carries process
startup, OPC unzip, composite build, serialisation and file I/O, and this box
has a measured ~6% run-to-run spread. The corpus runs were 390,116 ms on 2.5.1
against 366,810 / 374,190 / 368,737 on 2.6.0: consistent in direction, too
close to the noise floor to quote.

TWO MEASUREMENTS I FIRST REPORTED WERE WRONG, both because an uncommitted pin
change followed `git checkout` onto main -- an uncommitted edit to a file that
is identical on both branches does not get reverted by switching. So a "2.5.1
re-baseline" and a "2.5.1" transform run were both 2.6.0. The benchmark printed
the assembly version it had loaded, which is the only reason it surfaced; the
numbers above come from runs that assert the expected version is in bin before
measuring. For an A/B the pin belongs in a commit, not the working tree.

Also verified
  full suite              383 tests, 382 pass, 1 skip (unchanged)
  build                   clean, 0 warnings, 0 errors
  corpus output           13 of 13 byte-identical, across two runs
  restore --locked-mode   exit 0
  transitive              Core 2.0.0 -> 2.1.0, XQuery 2.5.1 -> 2.6.0

Byte-identical output alongside a speedup is the expected combination: every
perf commit changes how apply-templates walks children and what it allocates,
not what it emits.

The release also carries a security fix, bd57e34, so a cancelled transformation
stops inside XPath, regex matching and sorting. docmd passes
CancellationToken through ConvertAsync but has no caller that cancels today, so
this costs nothing now and is correct when batch mode arrives.

Core moving 2.0.0 -> 2.1.0 is transitive here and harmless. It is NOT harmless
for the phoenixml repo, which implements IContainer; whoever takes 2.6.0 there
should expect a breaking change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018wZEgtuzaZPswykiGEBxbz
Records the 14% transform gain with the measurement conditions, and why the
end-to-end figure is not quotable on this hardware.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018wZEgtuzaZPswykiGEBxbz
@elvogel
elvogel marked this pull request as draft October 7, 2026 05:56
@elvogel

elvogel commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

On hold — waiting for the next engine release.

2.6.0 is verified good: ~14% faster on the transform, 13 of 13 corpus documents byte-identical, 383 tests passing. But the next version carries further performance and memory-footprint work, and verification costs the same whether it crosses one release or two — so this waits and goes in as one bump rather than two.

Nothing here is wasted. When the next version lands, the comparison gets cheaper because the baselines now exist, measured on the same box with the same harness:

engine transform, 2,000 paragraphs (min of 7)
2.5.1 35,503 ms
2.6.0 30,387 ms

Two things to carry forward into that measurement:

  • Measure allocations as well as time. Four of 2.6.0's perf commits quote allocation reductions (−35%, −58%) rather than wall clock, and the next release is described as improving memory footprint. Time alone would under-report it.
  • The A/B harness needs the pin in a commit, not the working tree. An uncommitted pin edit followed git checkout onto main here and silently produced two wrong readings before the benchmark's own assembly-version line caught it.

Converted to draft rather than closed so the branch, the measurements and this reasoning stay in one place.

@elvogel

elvogel commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Target: engine 2.7.0. Not yet published as of 2026-10-07.

This draft holds the 2.6.0 bump. When 2.7.0 publishes, it supersedes this PR. 2.7.0 carries further performance and memory work, so one bump replaces two.

Protocol for that test, fixed in advance:

  1. Commit the pin on a branch. Do not measure from a dirty working tree.
  2. Assert the expected assembly version in bin before each run.
  3. Time MarkdownTransform.RunAsync in-process. CLI timing cannot resolve the effect.
  4. Measure allocations as well as time.
  5. Compare against both baselines below.
engine transform, 2,000 paragraphs, min of 7
2.5.1 35,503 ms
2.6.0 30,387 ms

@elvogel

elvogel commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Superseded by #49, which takes 2.7.0 instead.

2.6.0 leaves GHSA-xxjq-rwpx-m5ww open (it names everything up to and including 2.6.0), so this
branch cannot be the one that ships. Closing.

One thing measured here is worth carrying forward rather than discarding: 2.6.0 is the fastest
of the three.
Corpus transform totals, minimum of 5 runs per document, three isolated
worktrees — 2.5.1 212,828 ms, 2.6.0 194,280 ms, 2.7.0 204,756 ms. So 2.7.0 gives back about two
thirds of 2.6.0's gain. Filed upstream as phoenixmldb-xslt#319; if that is fixed, the gain
returns without giving up the advisory fix.

The release this branch was held for never arrived: nothing between the 2.6.0 and 2.7.0 tags is
performance work.

@elvogel elvogel closed this Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant