Skip to content

fix: an interrupted ordered list continues its numbering (#39) - #46

Merged
elvogel merged 1 commit into
mainfrom
fix/39-list-start
Oct 4, 2026
Merged

elvogel merged 1 commit into
mainfrom
fix/39-list-start

Conversation

@elvogel

@elvogel elvogel commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Closes #39.

Word stores no list object: every item is a top-level w:p carrying a numId. w:body groups
adjacent paragraphs sharing one, so a note in the middle of a procedure splits one list into
two runs and the second began again at 1. Word shows it continuing, because the paragraphs share
a numId.

md:list/@start already existed in the vocabulary and the serialiser already read it —
MarkdownSerializer line 274, ReadInt(list, "start") ?? 1. Only the stylesheet never emitted
it, exactly as docs/deferred-work.md said.

Scope

Level 0 only, and only emitted when the count is non-zero:

  • Word restarts a nested level under each new parent item by default, so a sub-list beginning
    at 1 is already correct. Counting preceding items at every level would make the second parent's
    sub-list continue at c. — turning a fix into a new bug. Two counterpart tests cover that
    (TwoSeparateOrderedLists_EachStartAtOne, AnInterruptedBulletList_IsUnaffected); both passed
    before the change and are regression guards.
  • Emitting nothing for the uninterrupted case keeps every document without this problem byte for
    byte what it was.
  • One predicate with and rather than two chained — chained predicates are quadratic on this
    engine (phoenixmldb-xslt#10, #95).

What the corpus proves about this: nothing

The reference corpus contains zero w:numPr paragraphs. These WordPerfect conversions number
manually, with literal text and tabs, so build-list never runs on them. My first two
measurements — "0 of 13 documents changed" and "386,170 ms, in band" — were both measuring a code
path that never executed. Stated plainly because a vacuous pass reads exactly like a clean one.

So the cost was measured against a scaling probe

Worst case for a preceding-sibling scan: one numId split into as many runs as possible.

Correctness holds at scale — 1,600 items continuing across 800 interruptions, final marker
1600.

runs main (control) with the fix delta
50 10,356 ms 10,472 ms +1.1%
200 25,717 ms 25,589 ms −0.5%
800 245,664 ms 237,050 ms −3.5%

The superlinear curve belongs to main: 4× the input costs 9.5× the time at the 200→800 step
without this change, independently reproducing what docs/limitations.md already records about
cost above ~4,000 paragraphs. The scan is O(runs²) in theory and invisible in practice next to
that.

Three data points rather than two on purpose: 50→200 grew only 2.4× for 4× input, which in
isolation looks better than linear. That is startup cost dominating a small document, and a
two-point measurement here would have been reassuring and wrong.

Verified

372 tests, 371 pass, 1 skip.

🤖 Generated with Claude Code

https://claude.ai/code/session_018wZEgtuzaZPswykiGEBxbz

Word stores no list object: every item is a top-level w:p carrying a numId.
w:body groups ADJACENT paragraphs sharing one, so a note in the middle of a
procedure splits one list into two runs and the second began again at 1. Word
shows it continuing, because the paragraphs share a numId.

md:list/@start already existed in the vocabulary and the serialiser already
read it (MarkdownSerializer line 274, ReadInt(list, "start") ?? 1). Only the
stylesheet never emitted it, which is exactly what docs/deferred-work.md said.

Scoped to level 0, and only emitted when the count is non-zero:

  - Word restarts a nested level under each new parent item by default, so a
    sub-list beginning at 1 is already correct. Counting preceding items at
    every level would make the second parent's sub-list continue at "c.",
    turning a fix into a new bug. The two counterpart tests cover that.
  - Emitting nothing for the ordinary uninterrupted case keeps every document
    without this problem byte for byte what it was.
  - One predicate with `and` rather than two chained: chained predicates are
    quadratic on this engine (phoenixmldb-xslt#10 and #95).

WHAT THE CORPUS PROVES ABOUT THIS: NOTHING. The reference corpus contains ZERO
w:numPr paragraphs -- these WordPerfect conversions number manually, with
literal text and tabs -- so build-list never runs on them. "0 of 13 documents
changed" and "386,170 ms, in band" were both measuring a code path that never
executed. Recorded because a vacuous pass reads exactly like a clean one.

So the scan's cost was measured against a synthetic scaling probe instead: one
numId split into as many runs as possible, which is the worst case for a
preceding-sibling scan. Correctness holds at scale -- 1,600 items continuing
across 800 interruptions, last marker "1600." -- and the cost is not mine:

  runs   main (control)   with the fix   delta
    50        10,356 ms      10,472 ms   +1.1%
   200        25,717 ms      25,589 ms   -0.5%
   800       245,664 ms     237,050 ms   -3.5%

The superlinear curve belongs to main: 4x the input costs 9.5x the time at the
200-to-800 step WITHOUT this change, which independently reproduces what
docs/limitations.md already records about cost above ~4,000 paragraphs. The
scan is O(runs squared) in theory and invisible in practice next to that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018wZEgtuzaZPswykiGEBxbz
@elvogel
elvogel force-pushed the fix/39-list-start branch from b381afd to c225d2d Compare October 4, 2026 02:42
@elvogel
elvogel merged commit 7b96a27 into main Oct 4, 2026
1 check passed
@elvogel
elvogel deleted the fix/39-list-start branch October 4, 2026 02:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

An interrupted ordered list restarts at 1. instead of continuing

1 participant