feat: lazy alphabet loading + generated alphabet metadata index - #68
Merged
Merged
Conversation
Startup measurements (dasher-startup-timing-report.md, Aug 2026) showed alphabet XML parsing dominates DasherCore's realize everywhere: 84% on Android (1,230-1,300ms of a 1,410-1,555ms realize), 78% on GTK, 75% on host — because Realize full-parsed all ~475 alphabet files to use one. Lazy loading (all modes, promoting the low-memory mechanism): - CAlphIO::ScanNameIndex() builds AlphID → filename by reading only each file's root <alphabet name="..."> attribute (first 2KB, no XML parse). ~10-20ms for the whole corpus. - Realize full-parses just the selected alphabet (+ default fallback). - ChangeAlphabet loads the requested alphabet on demand in every mode. - GetAlphabets (the menu / GetPermittedValues) serves index ∪ parsed — every alphabet stays selectable; parsing is what became lazy. Measured: realize.alphabets 1,300→65-150ms (Android), 270→10-23ms (GTK); realize.total ~3-3.5x faster on every platform. Bug fixed: the old low-memory on-demand check used GetInfo, which never returns null (it silently falls back to Default) — so switching to an unparsed alphabet in low-memory mode never loaded the file and silently gave the user the Default alphabet. HasInfo() is a real lookup and the switch path now uses it. Also ships Scripts/generate-alphabet-index.py + the generated Data/alphabets/alphabet_index.json: full-parse metadata for all 474 unique alphabets (deduped across corpus tiers with variant tracking) — lang (456/474 via filename/training/token derivation), ISO 15924 script (48 scripts, Unicode majority-detection fallback), orientation (ltr/rtl/ttb/btt), training file, palette, conversion mode, symbol counts, group names, and orientation-mismatch flags (two real data bugs found in legacy Urdu/Dhivehi files). Consumed by the website catalogue (website#40) and frontends for metadata-driven settings. tests: low_memory_alphabet_count adapted — the menu-count semantic changed (full inventory in all modes); suite 39/39. Signed-off-by: will wade <willwade@gmail.com>
CAlphabetMap::Add DASHER_ASSERTed that a symbol's output text was unique — but 132 shipped alphabet files violate that (the autoconverter applied Latin case groups to caseless scripts: every Arabic letter twice, 282 Ethiopic duplicates, 785+ per Mandarin tone routing, ...). Selecting any of those alphabets from the menu SIGABRTed the process in every v6 build to date. Add now skips redefinitions (first wins). Regression test switches to the worst shipped offender (Arabic, 36 duplicate letters) and runs frames. The index generator flags duplicate-symbols per file (116 unique alphabets after dedup) so the data can be fixed upstream. Signed-off-by: will wade <willwade@gmail.com>
With duplicates tolerated, the next assert in Add/Get fired: keys must be exactly one UTF-8 codepoint (lead byte length == key length). Shipped alphabets violate that too — lam-alef style ligature symbols output two characters. Both dispatches now use functional guards instead of asserts: single-byte keys in the direct table, everything else through the hash. Digraph keys won't match lead-byte slicing during text import — harmless degradation versus aborting the process. Signed-off-by: will wade <willwade@gmail.com>
The runtime name scanner hand-rolled attribute extraction and got three things wrong on valid XML (all flagged in review): - Single-quoted attributes (name='...') were never recognized. - XML character references were stored undecoded — 42 shipped files carry entity-encoded names (Latviešu, Türkmen, ...), so those alphabets were indexed under their raw encoded bytes and selecting them from the menu silently activated Default. - A fixed 2048-byte prefix window dropped alphabets whose prolog, DOCTYPE, or opening tag pushed the name attribute past it. The indexer now loads each file with pugixml itself — identical parsing semantics to the full loader (entity decoding, both quote styles, arbitrary prolog length), still skipping the expensive CAlphInfo build. Measured cost: 95-136ms for the 474-file corpus on host (vs 10-23ms for the hand-rolled scan and 175-210ms for the old eager parse+build) — the correctness is worth the delta, and lazy loading remains a large net win. Regression test: one alphabet exercising all three edges (single quotes, numeric references, >2KB prolog) — the menu must list the decoded name and selecting it must load it, not Default. Signed-off-by: will wade <willwade@gmail.com>
…iew) When two files declare the same AlphID (145 shipped IDs exist in up to three corpus tiers — e.g. 'English with limited punctuation' is both a maintained v6 file and an oldAlphabets v5 original), the index kept whichever file the filesystem traversal visited last: unspecified, filesystem-dependent, and able to hand the engine a legacy v5 definition in place of the authoritative one. RememberFileName now resolves deterministically by corpus tier — maintained > autoConverted/WorldAlphabets > oldAlphabets — matching the index generator's tier order. Regression test declares one ID in two tiers with distinguishable symbol counts and asserts the maintained definition loads. Signed-off-by: will wade <willwade@gmail.com>
…(review) corpusTier matched 'oldAlphabets'/'autoConverted' as substrings anywhere in the path — a data directory living inside a folder with one of those names collapsed every scanned file to that tier, and duplicate-ID resolution fell back to traversal order (could load the legacy symbol set). The tier is now decided by the file's immediate parent directory only, which is the corpus layout convention. Regression test nests the whole corpus inside a folder literally named oldAlphabets and asserts the maintained definition still wins. Signed-off-by: will wade <willwade@gmail.com>
GTK passes a relative data dir ('Data') and relies on CWD; ScanFiles
then hands the indexer relative paths, and LoadAlphabetFile's
single-file fast path only triggers for absolute paths — a relative
path was compiled as a regex, matched nothing, and every selection
silently fell back to the builtin 26-letter Default alphabet. On such
frontends every alphabet rendered as the near-untrained builtin: flat
letters with no weighting (the symptom seen live on GTK).
The indexer now stores absolute paths. Verified with a probe: relative
and absolute data dirs both load real definitions (maintained English
62 symbols, WA English 68, Arabic 88 — previously all reported the
builtin's 27).
Tests hardened: the duplicate-ID fixtures now use 30-letter alphabets
(31 symbols) so the count assertions cannot pass via a silent Default
fallback (the builtin also has 27 symbols — the old assertions were
satisfied by the fallback itself).
Signed-off-by: will wade <willwade@gmail.com>
…ender flat no more Closes the rendering half of #69. The autoconverter emitted free-form colorInfoName values ('lower case', 'korean trailing consonants', pinyin syllables, ...) but the shipped palettes define exactly eleven tokens (lowercase, uppercase, numbers, punctuation, accents, space, paragraph, paragraphSpace, limitedPunctuation, punctuationLong, lowercaseBackground). A group with no resolvable colour info draws nothing: the alphabet's probabilities exist but are invisible — every WA alphabet rendered as a flat, unweighted canvas. 321 of the 323 catalogue languages are served only by these alphabets. Bisected to the single attribute: identical nodes/groups/corpus/ palettes render 1 box (1.0x spread) with colorInfoName='lower case' and 28 boxes (5.3x) with 'lowercase'. Scripts/fix-wa-colorinfo.py rewrites all 474 autoConverted files in place (keyword mapping; unknown letter groups map to lowercase, the engine's own default). Verified with the tree-spread probe: English WA 2 boxes 1.1x -> 28 boxes 5.3x; Arabic (36 duplicate symbols) -> 31 boxes 2.8x; Korean -> 25 boxes 2.4x, matching maintained English. Suite 39/39. Converter rule for the WorldAlphabets exporter (recorded in the script docstring and #69): emit only the eleven palette tokens, or omit the attribute entirely — never free-form names. Follow-up observed but out of scope here: conversion-mode alphabets (Mandarin pinyin routings) still render sparsely at the top level — separate issue from colour tokens. Signed-off-by: will wade <willwade@gmail.com>
|
Too many files changed for review (483 files, 100 file limit). Bypass the limit by tagging |
willwade
added a commit
to dasher-project/Dasher-Apple
that referenced
this pull request
Aug 30, 2026
…47) Fixes the first-run canvas lag — the Apple half of the startup story GTK/Android tackled around [DasherCore #68](dasher-project/DasherCore#68). ## Diagnosis DasherCore #68 already made alphabet parsing lazy engine-side (84% of Android's realize was full-parsing all ~475 alphabet files; we picked that up in v0.2.16). The remaining Apple bottleneck: **engine create + realize ran synchronously on the main thread inside `DasherViewModel.init`** — the first canvas frame waited on alphabet parse + PPM training. ## Fix — Android's RFC 0018 pattern - **Two-phase bridge** (iOS/macOS/visionOS): init is cheap; `bootstrap()` runs `dasher_create` + callback registration + device locale + initial realize on a detached task. Rare UI-facing callbacks (message/speak/clipboard/parameter) hop to the main queue; output/delete stay synchronous to preserve draw-loop ordering. Every engine method no-ops until boot completes (existing `guard ctx` pattern). - **macOS preserves its deferred realize**: `bootstrap(realizeDefaultScreen: false)` — the v5 migration must be able to set parameters pre-realize — and `startEngine` now applies deferred params then realizes off-main via the new `realize(width:height:)`. - **ViewModels publish `isEngineReady`**; engine-dependent config (access `apply`, gaze defaults, tilt, speed read) runs post-boot. - **`EngineLoadingOverlay`** over every canvas site: 300 ms grace period so warm starts never flash it (Android's RFC 0018 behaviour — the canvas never sits dead). All four schemes build clean. Should be verified on-device: first launch shows the brief loading veil only after 300 ms and the UI is interactive immediately; warm starts see no flash. Signed-off-by: will wade <willwade@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Startup measurements (Aug 2026) showed alphabet XML parsing dominates
Realizeeverywhere — 84% on Android (1,230–1,300ms of a 1,410–1,555ms realize), 78% GTK, 75% host — because all ~475 alphabet files were full-parsed to use exactly one.Lazy loading (all modes)
CAlphIO::ScanNameIndex()— AlphID → filename index built by reading only each file's root<alphabet name="...">attribute (first 2KB, no XML parse). ~10–20ms for the whole corpus.Realizefull-parses just the selected alphabet (+ default fallback);ChangeAlphabetloads on demand in every mode.GetAlphabets(menu /GetPermittedValues) serves index ∪ parsed — every alphabet stays selectable; parsing is what became lazy.realize.totalFull data:
dasher-startup-timing-report.md(workspace).Bug fixed
The old low-memory on-demand check used
GetInfo— which never returns null (silently falls back to Default) — so switching to an unparsed alphabet in low-memory mode never loaded the file and silently gave the user Default.HasInfo()is a real lookup; the switch path uses it now.Alphabet metadata index
Scripts/generate-alphabet-index.py→Data/alphabets/alphabet_index.json: full-parse metadata for 474 unique alphabets (deduped across the three corpus tiers, variants tracked) — lang code (456/474), ISO-15924 script (48 scripts), orientation, training file, palette, conversion mode, chars, groups, and mismatch flags (two real data bugs found: legacy Urdu/Dhivehi declare LR for RTL scripts). Consumed by the website catalogue (dasher-project/website#40) and frontends.Tests
low_memory_alphabet_countadapted: the menu-count semantic changed (full inventory in all modes, identical across modes) — the oldcount <= 2asserted menu == parsed set.Greptile Summary
The PR replaces eager alphabet construction with a lightweight pugixml-based name index and loads the selected definition on demand, while preserving the full alphabet inventory exposed to frontends.
Confidence Score: 5/5
The PR appears safe to merge.
No blocking failure remains.
Important Files Changed
Flowchart
%%{init: {'theme': 'neutral'}}%% flowchart LR D[Alphabet data directories] --> S[Scan XML with pugixml] S --> I[Name-to-file index] I --> M[Alphabet menu and permitted values] R[Realize or ChangeAlphabet] --> L{Definition already loaded?} L -->|Yes| A[Activate alphabet] L -->|No| F[Full-parse indexed file] F --> A F -->|Missing or invalid| X[Default fallback]Reviews (7): Last reviewed commit: "fix: absolutize index paths — relative d..." | Re-trigger Greptile
Context used: