Repository navigation
fix: an interrupted ordered list continues its numbering (#39) - #46
Merged
Merged
Conversation
Word stores no list object: every item is a top-level w:p carrying a numId. w:body groups ADJACENT paragraphs sharing one, so a note in the middle of a procedure splits one list into two runs and the second began again at 1. Word shows it continuing, because the paragraphs share a numId. md:list/@start already existed in the vocabulary and the serialiser already read it (MarkdownSerializer line 274, ReadInt(list, "start") ?? 1). Only the stylesheet never emitted it, which is exactly what docs/deferred-work.md said. Scoped to level 0, and only emitted when the count is non-zero: - Word restarts a nested level under each new parent item by default, so a sub-list beginning at 1 is already correct. Counting preceding items at every level would make the second parent's sub-list continue at "c.", turning a fix into a new bug. The two counterpart tests cover that. - Emitting nothing for the ordinary uninterrupted case keeps every document without this problem byte for byte what it was. - One predicate with `and` rather than two chained: chained predicates are quadratic on this engine (phoenixmldb-xslt#10 and #95). WHAT THE CORPUS PROVES ABOUT THIS: NOTHING. The reference corpus contains ZERO w:numPr paragraphs -- these WordPerfect conversions number manually, with literal text and tabs -- so build-list never runs on them. "0 of 13 documents changed" and "386,170 ms, in band" were both measuring a code path that never executed. Recorded because a vacuous pass reads exactly like a clean one. So the scan's cost was measured against a synthetic scaling probe instead: one numId split into as many runs as possible, which is the worst case for a preceding-sibling scan. Correctness holds at scale -- 1,600 items continuing across 800 interruptions, last marker "1600." -- and the cost is not mine: runs main (control) with the fix delta 50 10,356 ms 10,472 ms +1.1% 200 25,717 ms 25,589 ms -0.5% 800 245,664 ms 237,050 ms -3.5% The superlinear curve belongs to main: 4x the input costs 9.5x the time at the 200-to-800 step WITHOUT this change, which independently reproduces what docs/limitations.md already records about cost above ~4,000 paragraphs. The scan is O(runs squared) in theory and invisible in practice next to that. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018wZEgtuzaZPswykiGEBxbz
elvogel
force-pushed
the
fix/39-list-start
branch
from
October 4, 2026 02:42
b381afd to
c225d2d
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #39.
Word stores no list object: every item is a top-level
w:pcarrying anumId.w:bodygroupsadjacent paragraphs sharing one, so a note in the middle of a procedure splits one list into
two runs and the second began again at
1.Word shows it continuing, because the paragraphs sharea
numId.md:list/@startalready existed in the vocabulary and the serialiser already read it —MarkdownSerializerline 274,ReadInt(list, "start") ?? 1. Only the stylesheet never emittedit, exactly as
docs/deferred-work.mdsaid.Scope
Level 0 only, and only emitted when the count is non-zero:
at 1 is already correct. Counting preceding items at every level would make the second parent's
sub-list continue at
c.— turning a fix into a new bug. Two counterpart tests cover that(
TwoSeparateOrderedLists_EachStartAtOne,AnInterruptedBulletList_IsUnaffected); both passedbefore the change and are regression guards.
byte what it was.
andrather than two chained — chained predicates are quadratic on thisengine (
phoenixmldb-xslt#10,#95).What the corpus proves about this: nothing
The reference corpus contains zero
w:numPrparagraphs. These WordPerfect conversions numbermanually, with literal text and tabs, so
build-listnever runs on them. My first twomeasurements — "0 of 13 documents changed" and "386,170 ms, in band" — were both measuring a code
path that never executed. Stated plainly because a vacuous pass reads exactly like a clean one.
So the cost was measured against a scaling probe
Worst case for a
preceding-siblingscan: onenumIdsplit into as many runs as possible.Correctness holds at scale — 1,600 items continuing across 800 interruptions, final marker
1600.main(control)The superlinear curve belongs to
main: 4× the input costs 9.5× the time at the 200→800 stepwithout this change, independently reproducing what
docs/limitations.mdalready records aboutcost above ~4,000 paragraphs. The scan is O(runs²) in theory and invisible in practice next to
that.
Three data points rather than two on purpose: 50→200 grew only 2.4× for 4× input, which in
isolation looks better than linear. That is startup cost dominating a small document, and a
two-point measurement here would have been reassuring and wrong.
Verified
372 tests, 371 pass, 1 skip.
🤖 Generated with Claude Code
https://claude.ai/code/session_018wZEgtuzaZPswykiGEBxbz