Skip to content

feat: lazy alphabet loading + generated alphabet metadata index - #68

Merged
willwade merged 8 commits into
mainfrom
feat/lazy-alphabets
Aug 28, 2026
Merged

willwade merged 8 commits into
mainfrom
feat/lazy-alphabets

Conversation

@willwade

@willwade willwade commented Aug 28, 2026 •

Copy link
Copy Markdown

Startup measurements (Aug 2026) showed alphabet XML parsing dominates Realize everywhere — 84% on Android (1,230–1,300ms of a 1,410–1,555ms realize), 78% GTK, 75% host — because all ~475 alphabet files were full-parsed to use exactly one.

Lazy loading (all modes)

  • CAlphIO::ScanNameIndex() — AlphID → filename index built by reading only each file's root <alphabet name="..."> attribute (first 2KB, no XML parse). ~10–20ms for the whole corpus.
  • Realize full-parses just the selected alphabet (+ default fallback); ChangeAlphabet loads on demand in every mode.
  • GetAlphabets (menu / GetPermittedValues) serves index ∪ parsed — every alphabet stays selectable; parsing is what became lazy.
realize.total Before After
Android app (emulator) 1,410–1,555ms 375–655ms
GTK 340–441ms 98–204ms
Host 235ms 93ms

Full data: dasher-startup-timing-report.md (workspace).

Bug fixed

The old low-memory on-demand check used GetInfo — which never returns null (silently falls back to Default) — so switching to an unparsed alphabet in low-memory mode never loaded the file and silently gave the user Default. HasInfo() is a real lookup; the switch path uses it now.

Alphabet metadata index

Scripts/generate-alphabet-index.py → Data/alphabets/alphabet_index.json: full-parse metadata for 474 unique alphabets (deduped across the three corpus tiers, variants tracked) — lang code (456/474), ISO-15924 script (48 scripts), orientation, training file, palette, conversion mode, chars, groups, and mismatch flags (two real data bugs found: legacy Urdu/Dhivehi declare LR for RTL scripts). Consumed by the website catalogue (dasher-project/website#40) and frontends.

Tests

  • low_memory_alphabet_count adapted: the menu-count semantic changed (full inventory in all modes, identical across modes) — the old count <= 2 asserted menu == parsed set.
  • Suite 39/39 green.

Greptile Summary

The PR replaces eager alphabet construction with a lightweight pugixml-based name index and loads the selected definition on demand, while preserving the full alphabet inventory exposed to frontends.

  • Adds deterministic maintained/converted/legacy precedence for duplicate alphabet IDs.
  • Adds a generated metadata catalogue for the packaged alphabet corpus.
  • Updates low-memory and scanner regression coverage for lazy discovery and switching.

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Important Files Changed

Filename Overview
src/DasherCore/Alphabet/AlphIO.cpp Implements pugixml-based alphabet discovery, indexed filenames, lazy full parsing, and deterministic cross-tier duplicate resolution; the displayed scanner and precedence regressions are addressed.
src/DasherCore/Alphabet/AlphIO.h Extends the alphabet loader interface and state for indexed discovery and explicit loaded-definition checks.
src/DasherCore/DasherInterfaceBase.cpp Integrates indexed discovery and on-demand parsing into realization and alphabet changes.
src/DasherCore/Alphabet/AlphabetMap.cpp Adjusts alphabet map behavior to support the revised alphabet-loading lifecycle without an accepted defect.
src/DasherCore/AlphabetManager.cpp Updates manager behavior for the lazy-loaded alphabet lifecycle without an accepted defect.
Scripts/generate-alphabet-index.py Generates the committed alphabet metadata catalogue from repository-controlled XML definitions.
Data/alphabets/alphabet_index.json Adds generated metadata for 474 unique packaged alphabet definitions.
tests/test_low_memory.cpp Updates inventory expectations and covers scanner parity, lazy selection, and duplicate-tier precedence regressions.

Flowchart

%%{init: {'theme': 'neutral'}}%%
flowchart LR
  D[Alphabet data directories] --> S[Scan XML with pugixml]
  S --> I[Name-to-file index]
  I --> M[Alphabet menu and permitted values]
  R[Realize or ChangeAlphabet] --> L{Definition already loaded?}
  L -->|Yes| A[Activate alphabet]
  L -->|No| F[Full-parse indexed file]
  F --> A
  F -->|Missing or invalid| X[Default fallback]
Loading

Reviews (7): Last reviewed commit: "fix: absolutize index paths — relative d..." | Re-trigger Greptile

Context used:

Startup measurements (dasher-startup-timing-report.md, Aug 2026) showed
alphabet XML parsing dominates DasherCore's realize everywhere: 84% on
Android (1,230-1,300ms of a 1,410-1,555ms realize), 78% on GTK, 75% on
host — because Realize full-parsed all ~475 alphabet files to use one.

Lazy loading (all modes, promoting the low-memory mechanism):
- CAlphIO::ScanNameIndex() builds AlphID → filename by reading only each
  file's root <alphabet name="..."> attribute (first 2KB, no XML
  parse). ~10-20ms for the whole corpus.
- Realize full-parses just the selected alphabet (+ default fallback).
- ChangeAlphabet loads the requested alphabet on demand in every mode.
- GetAlphabets (the menu / GetPermittedValues) serves index ∪ parsed —
  every alphabet stays selectable; parsing is what became lazy.
  Measured: realize.alphabets 1,300→65-150ms (Android), 270→10-23ms
  (GTK); realize.total ~3-3.5x faster on every platform.

Bug fixed: the old low-memory on-demand check used GetInfo, which never
returns null (it silently falls back to Default) — so switching to an
unparsed alphabet in low-memory mode never loaded the file and silently
gave the user the Default alphabet. HasInfo() is a real lookup and the
switch path now uses it.

Also ships Scripts/generate-alphabet-index.py + the generated
Data/alphabets/alphabet_index.json: full-parse metadata for all 474
unique alphabets (deduped across corpus tiers with variant tracking) —
lang (456/474 via filename/training/token derivation), ISO 15924 script
(48 scripts, Unicode majority-detection fallback), orientation
(ltr/rtl/ttb/btt), training file, palette, conversion mode, symbol
counts, group names, and orientation-mismatch flags (two real data bugs
found in legacy Urdu/Dhivehi files). Consumed by the website catalogue
(website#40) and frontends for metadata-driven settings.

tests: low_memory_alphabet_count adapted — the menu-count semantic
changed (full inventory in all modes); suite 39/39.

Signed-off-by: will wade <willwade@gmail.com>
Comment thread src/DasherCore/Alphabet/AlphIO.cpp Outdated
CAlphabetMap::Add DASHER_ASSERTed that a symbol's output text was unique
— but 132 shipped alphabet files violate that (the autoconverter applied
Latin case groups to caseless scripts: every Arabic letter twice, 282
Ethiopic duplicates, 785+ per Mandarin tone routing, ...). Selecting any
of those alphabets from the menu SIGABRTed the process in every v6
build to date.

Add now skips redefinitions (first wins). Regression test switches to
the worst shipped offender (Arabic, 36 duplicate letters) and runs
frames. The index generator flags duplicate-symbols per file (116
unique alphabets after dedup) so the data can be fixed upstream.

Signed-off-by: will wade <willwade@gmail.com>
Comment thread src/DasherCore/Alphabet/AlphIO.cpp Outdated
With duplicates tolerated, the next assert in Add/Get fired: keys must
be exactly one UTF-8 codepoint (lead byte length == key length). Shipped
alphabets violate that too — lam-alef style ligature symbols output two
characters. Both dispatches now use functional guards instead of
asserts: single-byte keys in the direct table, everything else through
the hash. Digraph keys won't match lead-byte slicing during text
import — harmless degradation versus aborting the process.

Signed-off-by: will wade <willwade@gmail.com>
Comment thread src/DasherCore/Alphabet/AlphIO.cpp Outdated
The runtime name scanner hand-rolled attribute extraction and got three
things wrong on valid XML (all flagged in review):

- Single-quoted attributes (name='...') were never recognized.
- XML character references were stored undecoded — 42 shipped files
  carry entity-encoded names (Latvie&#353;u, T&#x00FC;rkmen, ...), so
  those alphabets were indexed under their raw encoded bytes and
  selecting them from the menu silently activated Default.
- A fixed 2048-byte prefix window dropped alphabets whose prolog,
  DOCTYPE, or opening tag pushed the name attribute past it.

The indexer now loads each file with pugixml itself — identical parsing
semantics to the full loader (entity decoding, both quote styles,
arbitrary prolog length), still skipping the expensive CAlphInfo build.
Measured cost: 95-136ms for the 474-file corpus on host (vs 10-23ms for
the hand-rolled scan and 175-210ms for the old eager parse+build) — the
correctness is worth the delta, and lazy loading remains a large net
win.

Regression test: one alphabet exercising all three edges (single
quotes, numeric references, >2KB prolog) — the menu must list the
decoded name and selecting it must load it, not Default.

Signed-off-by: will wade <willwade@gmail.com>
Comment thread src/DasherCore/Alphabet/AlphIO.cpp Outdated
…iew)

When two files declare the same AlphID (145 shipped IDs exist in up to
three corpus tiers — e.g. 'English with limited punctuation' is both a
maintained v6 file and an oldAlphabets v5 original), the index kept
whichever file the filesystem traversal visited last: unspecified,
filesystem-dependent, and able to hand the engine a legacy v5
definition in place of the authoritative one.

RememberFileName now resolves deterministically by corpus tier —
maintained > autoConverted/WorldAlphabets > oldAlphabets — matching the
index generator's tier order. Regression test declares one ID in two
tiers with distinguishable symbol counts and asserts the maintained
definition loads.

Signed-off-by: will wade <willwade@gmail.com>
Comment thread src/DasherCore/Alphabet/AlphIO.cpp Outdated
…(review)

corpusTier matched 'oldAlphabets'/'autoConverted' as substrings anywhere
in the path — a data directory living inside a folder with one of those
names collapsed every scanned file to that tier, and duplicate-ID
resolution fell back to traversal order (could load the legacy symbol
set). The tier is now decided by the file's immediate parent directory
only, which is the corpus layout convention. Regression test nests the
whole corpus inside a folder literally named oldAlphabets and asserts
the maintained definition still wins.

Signed-off-by: will wade <willwade@gmail.com>
GTK passes a relative data dir ('Data') and relies on CWD; ScanFiles
then hands the indexer relative paths, and LoadAlphabetFile's
single-file fast path only triggers for absolute paths — a relative
path was compiled as a regex, matched nothing, and every selection
silently fell back to the builtin 26-letter Default alphabet. On such
frontends every alphabet rendered as the near-untrained builtin: flat
letters with no weighting (the symptom seen live on GTK).

The indexer now stores absolute paths. Verified with a probe: relative
and absolute data dirs both load real definitions (maintained English
62 symbols, WA English 68, Arabic 88 — previously all reported the
builtin's 27).

Tests hardened: the duplicate-ID fixtures now use 30-letter alphabets
(31 symbols) so the count assertions cannot pass via a silent Default
fallback (the builtin also has 27 symbols — the old assertions were
satisfied by the fallback itself).

Signed-off-by: will wade <willwade@gmail.com>
…ender flat no more

Closes the rendering half of #69. The autoconverter emitted free-form
colorInfoName values ('lower case', 'korean trailing consonants',
pinyin syllables, ...) but the shipped palettes define exactly eleven
tokens (lowercase, uppercase, numbers, punctuation, accents, space,
paragraph, paragraphSpace, limitedPunctuation, punctuationLong,
lowercaseBackground). A group with no resolvable colour info draws
nothing: the alphabet's probabilities exist but are invisible — every
WA alphabet rendered as a flat, unweighted canvas. 321 of the 323
catalogue languages are served only by these alphabets.

Bisected to the single attribute: identical nodes/groups/corpus/
palettes render 1 box (1.0x spread) with colorInfoName='lower case'
and 28 boxes (5.3x) with 'lowercase'.

Scripts/fix-wa-colorinfo.py rewrites all 474 autoConverted files in
place (keyword mapping; unknown letter groups map to lowercase, the
engine's own default). Verified with the tree-spread probe: English WA
2 boxes 1.1x -> 28 boxes 5.3x; Arabic (36 duplicate symbols) -> 31
boxes 2.8x; Korean -> 25 boxes 2.4x, matching maintained English.
Suite 39/39.

Converter rule for the WorldAlphabets exporter (recorded in the script
docstring and #69): emit only the eleven palette tokens, or omit the
attribute entirely — never free-form names.

Follow-up observed but out of scope here: conversion-mode alphabets
(Mandarin pinyin routings) still render sparsely at the top level —
separate issue from colour tokens.

Signed-off-by: will wade <willwade@gmail.com>
@greptile-apps

greptile-apps Bot commented Aug 28, 2026

Copy link
Copy Markdown

Too many files changed for review (483 files, 100 file limit).

Bypass the limit by tagging @greptile-apps to review.

@willwade
willwade merged commit 41e1c9a into main Aug 28, 2026
14 checks passed
@willwade
willwade deleted the feat/lazy-alphabets branch August 28, 2026 23:59
willwade added a commit to dasher-project/Dasher-Apple that referenced this pull request Aug 30, 2026
…47)

Fixes the first-run canvas lag — the Apple half of the startup story
GTK/Android tackled around [DasherCore
#68](dasher-project/DasherCore#68).

## Diagnosis
DasherCore #68 already made alphabet parsing lazy engine-side (84% of
Android's realize was full-parsing all ~475 alphabet files; we picked
that up in v0.2.16). The remaining Apple bottleneck: **engine create +
realize ran synchronously on the main thread inside
`DasherViewModel.init`** — the first canvas frame waited on alphabet
parse + PPM training.

## Fix — Android's RFC 0018 pattern
- **Two-phase bridge** (iOS/macOS/visionOS): init is cheap;
`bootstrap()` runs `dasher_create` + callback registration + device
locale + initial realize on a detached task. Rare UI-facing callbacks
(message/speak/clipboard/parameter) hop to the main queue; output/delete
stay synchronous to preserve draw-loop ordering. Every engine method
no-ops until boot completes (existing `guard ctx` pattern).
- **macOS preserves its deferred realize**:
`bootstrap(realizeDefaultScreen: false)` — the v5 migration must be able
to set parameters pre-realize — and `startEngine` now applies deferred
params then realizes off-main via the new `realize(width:height:)`.
- **ViewModels publish `isEngineReady`**; engine-dependent config
(access `apply`, gaze defaults, tilt, speed read) runs post-boot.
- **`EngineLoadingOverlay`** over every canvas site: 300 ms grace period
so warm starts never flash it (Android's RFC 0018 behaviour — the canvas
never sits dead).

All four schemes build clean. Should be verified on-device: first launch
shows the brief loading veil only after 300 ms and the UI is interactive
immediately; warm starts see no flash.

Signed-off-by: will wade <willwade@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant