Skip to content

feat(trainer): longest-match symbol lookup (RFC 0020 clause 4) - #98

Merged
willwade merged 4 commits into
mainfrom
feat/longest-match-trainer
Sep 22, 2026
Merged

willwade merged 4 commits into
mainfrom
feat/longest-match-trainer

Conversation

@willwade

@willwade willwade commented Sep 22, 2026 •

Copy link
Copy Markdown

Makes multi-codepoint symbols (digraphs, ZWJ families, VS16 skin tones) trainable: SymbolStream::next probes multi-codepoint map keys longest-first before the single-character path. First PR of RFC 0020.

  • CAlphabetMap::Add collects multi-codepoint keys (copies, sorted longest-first) + LongestMatch() probe + MaxKeyLen()
  • next() probes before single-char dispatch (❤️ beats its prefix ❤); ensureLookahead(want) extracted from findNext keeps MaxKeyLen bytes buffered across 1024-byte window refills
  • 6 unit tests + CAPI end-to-end (flat alphabet, corpus of only multi-cp tokens → top child 1.94× uniform; failed pre-fix)
  • Emoji corpus test relaxed to whole-node tokens per RFC
  • Full suite 47/47 locally

RetriggerConfidence Score: 4/5

The PR is not yet safe to merge because escaped adaptive-training contexts containing multi-codepoint symbols are reconstructed incorrectly.

Findings

  1. P1 Raw parsing corrupts context ▶
  2. P1 Longest match hides delimiter ▶

Summary

This PR adds longest-match training support for multi-codepoint alphabet symbols and protects structural annotation delimiters with raw, single-codepoint stream operations.

  • Registers multi-codepoint alphabet keys and probes them longest-first across stream refills.
  • Tracks the complete previously consumed token for peekBack().
  • Adds raw stream operations for conversion annotations and context-switch delimiters.
  • Adds internal, C API, refill-boundary, prefix, and delimiter regression coverage.
  • The raw mode is applied too broadly to context-switch body text, which prevents multi-codepoint symbols there from seeding the language-model context correctly.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
    A[Training stream] --> B{Context escape?}
    B -- No --> C[Longest-match next]
    C --> D[Learn alphabet symbol]
    B -- Yes --> E[Read delimiter as one raw codepoint]
    E --> F[Read saved document context]
    F --> G[Enter resolved symbols into model context]
    G --> H[Read closing delimiter structurally]
    H --> C
Loading

Reviews (4) · Last reviewed commit: "fix(trainer): greptile P1 r3 — raw mode ..."

SymbolStream::next read exactly one unicode character per lookup (a
2002-era design: 'we do not support multi-unicode-character symbols'),
so multi-codepoint symbols — digraph outputs, ZWJ emoji families, VS16
skin tones — could never be matched from training text. The trainer
split them into per-codepoint unknowns: 31 of the emoji alphabet's 308
nodes were permanently untrainable.

CAlphabetMap now collects multi-codepoint keys at Add time (copies —
the Entries vector reallocates) sorted longest-first, and next() probes
them before the single-character path, so ❤️ wins over its first-code-
point prefix ❤. findNext's refill logic is extracted to
ensureLookahead(want), reused to keep MaxKeyLen bytes buffered across
the 1024-byte window refills.

Tests: 6 unit cases (whole-sequence match, prefix preference, ASCII
digraphs, buffer-boundary straddle, empty probe list, unknown
degradation) + a CAPI end-to-end case (flat alphabet trained on only
multi-codepoint tokens: top child reaches 1.94x uniform mass; the
assertion that failed pre-fix). The emoji corpus test relaxes from
single-codepoint-only to whole-node tokens. Full suite: 47/47.

Signed-off-by: will wade <willwade@gmail.com>
Comment thread src/DasherCore/Alphabet/AlphabetMap.cpp
Comment thread src/DasherCore/Alphabet/AlphabetMap.cpp
Review loop 1 (7/10) findings:

- peekBack's contract broke under longest-match: the backward buffer
  walk returned only the FINAL codepoint of a multi-codepoint key, so
  an alphabet combining context-escape delimiters with a multi-char key
  ending in the delimiter char would falsely terminate a context block
  (latent — no shipped alphabet triggers it). next() now records the
  bytes it consumed and peekBack replays them; safe across window
  shifts by construction.
- Stale 'we do not support multi-unicode-character symbols' comments
  updated; >1024-byte-key limitation documented at Add.
- New tests: EOF-mid-key degradation, duplicate Add registers once,
  peekBack-after-match returns the whole 18-byte key.
- clang-format over touched files.

Full suite 47/47.

Signed-off-by: will wade <willwade@gmail.com>
… bound

- P1: annotation readers (Routing/Mandarin conversion trainers,
  CTrainer::readEscape) record the peeked token then advance via next();
  with a multi-codepoint key at the position, peekAhead returned only
  its first codepoint while next() consumed the whole key — the route/
  pronunciation was never learned. peekAhead now takes the map and
  probes longest-match, returning exactly the bytes the following
  next() consumes. All three call sites updated.
- P2: keys of STREAM_WINDOW (1024) bytes or more can never be fully
  buffered for probing — now excluded at registration with the limit
  documented (they were equally dead through the per-codepoint path).
- New test: peekAhead/next agreement across a multi-codepoint key.

Full suite 47/47.

Signed-off-by: will wade <willwade@gmail.com>
Comment thread src/DasherCore/Alphabet/AlphabetMap.cpp
If a conversion annotation's stop delimiter ('>' etc.) prefixes a
configured multi-codepoint key, longest-match swallows it into the key:
the annotation loop misses its terminator and absorbs the remainder of
the training file as annotation text.

Structural parsing is grammar, not symbol content: nextRaw /
peekAheadRaw read exactly one codepoint (the pre-RFC behaviour), and
the three structural readers — CTrainer::readEscape (escape delimiters
and the escaped-context loop), Routing and Mandarin annotation
accumulation — use them. Longest-match remains for symbol training,
where shadowing is the intended semantics.

Regression test contrasts both modes over the same '>x' bytes.

Full suite 47/47.

Signed-off-by: will wade <willwade@gmail.com>
Comment thread src/DasherCore/Trainer.cpp
@willwade
willwade merged commit 74814dc into main Sep 22, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant