From 50e3b7abb3f94c0028a4cfcb76c527c4439847b3 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Mon, 20 Jul 2026 21:06:56 -0400 Subject: [PATCH 01/21] Use a lot less memory, and give it back when the system asks Reported on a TCL Flip 2 as "the whole app is a bit slow" (#83). Measured on an M5 (2.9 GB, Android 13, standardDebug, median of 5 cold starts) the app held 831 MB at peak, sat at 421 MB idle, and gave back nothing at all when the OS asked it to shrink, because nothing in the tree implemented memory-pressure handling. The biggest win needs no device gate. The on-device speech model costs ~267 MB while loaded (~101 MB of weights plus ~146 MB of onnxruntime arena) and was kept for the whole process on the chance of a mic tap many users never make. It is now dropped after two minutes unused and rebuilt on next use, so the instant first tap that was asked for on 2026-07-10 is kept while the session-long hold is not. Device-verified: scudo:secondary 111 MB to 9 MB at the 120 s mark, model rebuilding correctly afterwards. Idle PSS 421 MB to 299 MB on a phone that is not low-RAM at all. Nothing released under pressure. There was no onTrimMemory, onLowMemory or ComponentCallbacks2 anywhere, so a TRIM_MEMORY_COMPLETE freed 0 KB and the system's only remaining option was to kill us. VelaApp now fans every trim out through a new MemoryPressure holder, and the things that actually hold memory register a release: the speech model, the neural voice, MapLibre's native tile and sprite caches (MapView.onLowMemory was never called), all five hidden WebViews, and the image cache. WhisperRecognizer had no release path at all, so even Remove-model left ~267 MB resident for the rest of the process. Nothing adapted to the device either. There was no isLowRamDevice branch and the image cache was a flat 48 MB whatever the phone. Constrained devices now skip the startup preload of the speech model, cap images at 16 MB, skip the speculative WebView warm on every search, and fetch 8 ambient POI category terms instead of 15 with a smaller result pool. Roomier phones keep their existing behaviour. The low-RAM POI subset deliberately keeps school and park: the ambient layer filter-hides the basemap OSM poi layers at z14+, so those two have no second source and a first 6-term subset made every park and school pin vanish. Caught by an A/B screenshot, not by any test. Also stop shipping x86 and x86_64 native libraries. No phone Vela targets can execute them and libmaplibre.so alone carried 23 MB of them into every install. armeabi-v7a stays, since 32-bit ARM keypad phones are real. Measured, main vs this, low-RAM path: peak 831 MB to 581 MB (-30%), post-trim 397 MB to 246 MB (-38%), native heap 223 MB to 95 MB (-57%), cold start 4811 ms to 4333 ms. On a normal-RAM device idle drops 29%, post-trim 28% and native heap 44%. APK 97.4 MB to 92.0 MB. Debug builds honour `setprop debug.vela.lowram true` so the low-RAM path can be exercised on a dev phone, where it is otherwise dead code (every device we own reports lowRam=false heapClassMb=256). AGENTS.md: document the seam and the measurement traps, and correct the memory rule, which said the OverpassTrafficSignals/OverpassPois stream-parse follow-up was pending. It has been done for some time; the remaining buffered hot reader is the Google ambient path, and chasing the stale line wasted a pass. --- AGENTS.md | 51 +++++++- app/build.gradle.kts | 12 +- app/src/main/java/app/vela/VelaApp.kt | 28 ++++- .../main/java/app/vela/ui/MemoryPressure.kt | 110 ++++++++++++++++++ .../main/java/app/vela/ui/map/MapViewModel.kt | 14 ++- .../main/java/app/vela/ui/map/VelaMapView.kt | 10 ++ .../main/java/app/vela/voice/PiperSynth.kt | 13 +++ .../java/app/vela/voice/WhisperRecognizer.kt | 103 +++++++++++++++- .../java/app/vela/web/WebDirectionsFetcher.kt | 20 +++- .../main/java/app/vela/web/WebPhotoFetcher.kt | 19 +++ .../app/vela/web/WebPopularTimesFetcher.kt | 22 +++- .../java/app/vela/web/WebReviewsFetcher.kt | 20 +++- .../app/vela/web/WebStopDeparturesFetcher.kt | 20 +++- .../java/app/vela/core/data/LowRamMode.kt | 17 +++ .../core/data/google/GoogleMapsDataSource.kt | 29 ++++- docs/FEATURES.md | 2 + 16 files changed, 458 insertions(+), 32 deletions(-) create mode 100644 app/src/main/java/app/vela/ui/MemoryPressure.kt create mode 100644 core/src/main/java/app/vela/core/data/LowRamMode.kt diff --git a/AGENTS.md b/AGENTS.md index 7ea33030..af526d06 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -627,9 +627,56 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): (raises the ceiling ~2x); don't remove it. (2) **Any Overpass / large-HTTP-body reader MUST stream-parse** - `Json.decodeFromStream(body.byteStream())` into a tiny `@Serializable` DTO, NEVER `resp.body.string()` + `parseToJsonElement` (that held ~5-10x the wire size in transient heap and - OOM'd mid-read - the Flock `out body` fetch per pan did this; fixed in `OverpassAlprCameras`, - `OverpassTrafficSignals`/`OverpassPois` are the same pattern + a pending follow-up). And NEVER lower a + OOM'd mid-read - the Flock `out body` fetch per pan did this). And NEVER lower a per-viewport Overpass fetch's min-zoom without shrinking the box. + **The Overpass follow-up is DONE (audited 2026-07-20):** `OverpassAlprCameras`, `OverpassTrafficSignals` + AND `OverpassPois` all `decodeFromStream` today. This line previously said the latter two were pending, + which sent a low-RAM investigation chasing already-fixed code. The remaining fully-buffered hot reader + is the GOOGLE ambient path, not Overpass: `GoogleMapsDataSource.get()` ends in `.string()`, and + `GoogleResponse.parse` then makes a `substring` copy plus a full `JsonElement` DOM - times a 15-term + fan-out per pan. That, not Overpass, is the ~180 MB/12 s. It cannot simply `decodeFromStream`: the + payload is a positional nameless array walked by `at(0,1,3)` paths, so there is no DTO to decode into. + The levers that DO move it are the term count and the `!7i` pool size. + +- **Low-RAM devices are a FIRST-CLASS target, and the app now adapts to them (issue #83, 2026-07-20).** + D-pad-first means feature-phone-first, and those phones are memory-poor as well as small. + - **`app/ui/MemoryPressure.kt` is the one seam.** `init()` from `VelaApp` classifies the device + (`ActivityManager.isLowRamDevice` OR heap class <= 127 MB) and `VelaApp.onTrimMemory` fans every + `TRIM_MEMORY_*` out to registered holders. **Anything that allocates something large or NATIVE + must register a release callback.** Registration, never a Hilt entry point: reaching a singleton + from a trim would CONSTRUCT it, so the trim would allocate the very thing it is freeing. + - `LowRamMode.enabled` (`:core`) is the `:core`-visible mirror, pushed in by `VelaApp` - same seam + as `CategoryFilter.enabled`, because `:core` must never read an `:app` holder. + - **Verify the low-RAM path or it ships unverified.** Every dev phone we own reports + `lowRam=false heapClassMb=256`, so those branches are dead code locally. Debug builds honour + `adb shell setprop debug.vela.lowram true` (then relaunch); `setprop debug.vela.lowram false` + restores real detection. NB `setprop ""` is a syntax error, not a reset. + - **Measuring: `am send-trim-memory` REFUSES background levels on a foreground process** + ("Unable to set a background trim level on a foreground process"). Press HOME first. A harness + that discards that stderr measures NOTHING and reports a clean baseline - this happened here and + produced a whole benchmark of void numbers before the error was noticed. Always check it. + - Measured on an M5 (2.9 GB, Android 13, standardDebug, median of 5 cold starts) main vs the fix, + low-RAM path: peak PSS 831 MB -> 581 MB (-30%), post-trim 397 MB -> 246 MB (-38%), native heap + 223 MB -> 95 MB (-57%), cold start 4811 ms -> 4333 ms. Idle PSS run-to-run variance is +-60 MB, + so single idle readings prove nothing; compare post-trim, which is paired within a run. + - **The ASR model is the single largest reclaimable allocation: ~267 MB PSS** (~101 MB of weights + in `scudo:secondary` plus ~146 MB of onnxruntime arena in `scudo:primary`). It releases on a + severe trim, is NOT warmed at startup on a low-RAM device, and - on EVERY device - is dropped + after `REAP_IDLE_MS` (120 s) unused and rebuilt on next use. `WhisperRecognizer.release()` + declines while a listen is in flight - freeing the native recognizer under a running decode is a + use-after-free that takes the process down rather than throwing. + - **Do not reach for a device gate when an IDLE gate will do.** The warm-at-startup behaviour was + first made low-RAM-conditional, which protected the instant-first-mic-tap UX on roomier phones + but left them holding 267 MB all session. Reaping on idle keeps that UX AND reclaims the memory + everywhere: device-verified on the 2.9 GB M5, `scudo:secondary` 111 MB -> 9 MB at the 120 s mark + with the model rebuilding on next use. Idle PSS 421 MB -> 299 MB (-29%) on a NON-low-RAM device. + Ask "can this be released when unused?" before "which devices should get less?". + - **`MapView.onLowMemory()` must be called.** MapLibre's tile/glyph/sprite caches are native and + that is the only way to shrink them; nothing called it before. + - **When you gate a category fan-out down, check what has no SECOND source.** The low-RAM ambient + subset keeps `school` and `park` on purpose: the ambient layer filter-hides the basemap OSM poi + layers at z14+, so those two would vanish entirely. A first 6-term subset did exactly that and + was caught by an A/B screenshot, not by any test. ## Layout diff --git a/app/build.gradle.kts b/app/build.gradle.kts index a03ccff7..8194c01a 100644 --- a/app/build.gradle.kts +++ b/app/build.gradle.kts @@ -217,12 +217,18 @@ android { resources { excludes += "/META-INF/{AL2.0,LGPL2.1}" } // The neural-TTS runtime (ONNX Runtime + sherpa-onnx, from the vendored AAR) ships its .so // for all 4 ABIs; Vela targets arm64 phones, so drop the other ABIs' copies - they'd add - // ~65 MB for no device we support. MapLibre and other libs stay multi-ABI (untouched). + // ~65 MB for no device we support. + // + // x86/x86_64 go for EVERY native lib, not just the TTS ones (issue #83). Those two ABIs + // exist for emulators; no phone Vela targets can execute them, and libmaplibre.so alone + // shipped 11.6 MB of x86_64 plus 11.6 MB of x86 in every install. armeabi-v7a is KEPT: + // it is a real 32-bit ARM target and dropping it would silently strand those phones, + // which is the opposite of this issue's goal. jniLibs { excludes += listOf( "**/armeabi-v7a/libonnxruntime.so", "**/armeabi-v7a/libsherpa-onnx*.so", - "**/x86/libonnxruntime.so", "**/x86/libsherpa-onnx*.so", - "**/x86_64/libonnxruntime.so", "**/x86_64/libsherpa-onnx*.so", + "**/x86/*.so", + "**/x86_64/*.so", ) } } diff --git a/app/src/main/java/app/vela/VelaApp.kt b/app/src/main/java/app/vela/VelaApp.kt index 8b5fc595..db0a2c9c 100644 --- a/app/src/main/java/app/vela/VelaApp.kt +++ b/app/src/main/java/app/vela/VelaApp.kt @@ -39,15 +39,34 @@ class VelaApp : Application(), coil.ImageLoaderFactory { * ~128 MB of decoded gallery bitmaps by design, which is most of the "rapid place churn * runs into the ceiling" OOM (issue #182; measured: 3 gallery-bearing places grew the live * Dalvik heap 14 -> 94 MB). 48 MB still holds a couple of screens of thumbnails + a hero - * or two; everything else re-decodes from Coil's disk cache, which is untouched. */ + * or two; everything else re-decodes from Coil's disk cache, which is untouched. + * + * The cap is now a function of the device instead of one constant: a 48 MB bitmap cache is + * reasonable on a 2-3 GB phone and absurd on a keypad phone whose whole heap class is 96 MB + * (issue #83). Low-RAM devices get 16 MB, which still covers a screen of result thumbnails. + * [MemoryPressure.init] must run before this, and does - onCreate inits it first. */ override fun newImageLoader(): coil.ImageLoader = coil.ImageLoader.Builder(this) .memoryCache { coil.memory.MemoryCache.Builder(this) - .maxSizeBytes(48 * 1024 * 1024) + .maxSizeBytes(if (app.vela.ui.MemoryPressure.lowRam) 16 * 1024 * 1024 else 48 * 1024 * 1024) .build() } .build() + /** + * Hand OS memory pressure to every holder that owns a large or native allocation (issue #83). + * Before this existed nothing in the app implemented ComponentCallbacks2, so a + * TRIM_MEMORY_COMPLETE released nothing at all and the OS had no option but to kill us. + * Coil's own cache is trimmed here; everything else releases through [MemoryPressure]. + */ + override fun onTrimMemory(level: Int) { + super.onTrimMemory(level) + app.vela.ui.MemoryPressure.dispatch(level) + if (app.vela.ui.MemoryPressure.isSevere(level)) { + runCatching { coil.Coil.imageLoader(this).memoryCache?.clear() } + } + } + /** Apply the persisted in-app language to the Application context too (no-op when following the * system), so `getString` from the ViewModel/nav-notification also localizes - resolved at launch * from the saved pref (an in-session change re-reads it on next launch). */ @@ -64,6 +83,11 @@ class VelaApp : Application(), coil.ImageLoaderFactory { Timber.plant(DiagTree(diag)) if (BuildConfig.DEBUG) Timber.plant(Timber.DebugTree()) + // Device memory class first: the Coil cap and the eager-warm decisions below both read it. + app.vela.ui.MemoryPressure.init(this) + // Push the device class down to :core, which cannot read an :app holder (same seam as + // CategoryFilter.enabled). Gates the ambient POI fan-out in GoogleMapsDataSource. + app.vela.core.data.LowRamMode.enabled = app.vela.ui.MemoryPressure.lowRam Units.init(this) AppTheme.init(this) AppLocale.init(this) // resolve the app language (system default) → drives the nav-text locale diff --git a/app/src/main/java/app/vela/ui/MemoryPressure.kt b/app/src/main/java/app/vela/ui/MemoryPressure.kt new file mode 100644 index 00000000..30e8a652 --- /dev/null +++ b/app/src/main/java/app/vela/ui/MemoryPressure.kt @@ -0,0 +1,110 @@ +package app.vela.ui + +import android.app.ActivityManager +import android.content.ComponentCallbacks2 +import android.content.Context +import java.util.concurrent.CopyOnWriteArrayList +import timber.log.Timber + +/** + * Process-wide memory-pressure fan-out, in the same shape as the other app-level holders + * (`TransitLayer`, `AppTheme`): `init()` from `VelaApp`, then anything holding a large or native + * allocation registers a release callback. + * + * Why registration and not a Hilt entry point: reaching `WhisperRecognizer`/`PiperSynth` from + * `onTrimMemory` through an EntryPoint would CONSTRUCT them if they had never been used, so a + * trim would allocate the very models it is trying to free. A holder registers only once it + * actually owns something worth releasing, so a trim can never create work. + * + * Measured on the M5 (2.9 GB, Android 13, standardDebug) before this existed: TRIM_MEMORY_COMPLETE + * released 0 KB, because nothing in the app implemented ComponentCallbacks2 at all. + */ +object MemoryPressure { + + /** A registered releaser. [level] is a `ComponentCallbacks2.TRIM_MEMORY_*` constant. */ + fun interface Listener { + fun release(level: Int) + } + + private val listeners = CopyOnWriteArrayList() + + /** + * True when the OS classes this device as low-RAM (`ActivityManager.isLowRamDevice`) OR its + * heap class is small enough that our normal budgets do not fit. The heap-class arm matters: + * plenty of cheap keypad phones do NOT set the low-RAM system property yet still hand out a + * 96 MB heap class, and those are exactly the phones this work is for. + */ + @Volatile var lowRam: Boolean = false + private set + + /** The device's normal (non-large) heap class in MB. 0 until [init]. */ + @Volatile var heapClassMb: Int = 0 + private set + + fun init(context: Context) { + val am = context.getSystemService(Context.ACTIVITY_SERVICE) as? ActivityManager + heapClassMb = am?.memoryClass ?: 0 + val forced = forcedLowRam() + lowRam = forced ?: ((am?.isLowRamDevice == true) || (heapClassMb in 1..127)) + Timber.i( + "MemoryPressure init lowRam=%b heapClassMb=%d forced=%s", + lowRam, heapClassMb, forced?.toString() ?: "no", + ) + } + + /** + * Debug-only override so the low-RAM path can be exercised on a normal dev phone: + * + * adb shell setprop debug.vela.lowram true # then relaunch the app + * adb shell setprop debug.vela.lowram "" # back to real detection + * + * Without this the low-RAM branches are dead code on every device we actually own (the M5 dev + * phone reports heapClassMb=256, lowRam=false), which means they would ship unverified. Returns + * null when unset or on a non-debug build, so release behaviour is untouched. + */ + private fun forcedLowRam(): Boolean? { + if (!app.vela.BuildConfig.DEBUG) return null + val v = runCatching { + @Suppress("PrivateApi") + val sp = Class.forName("android.os.SystemProperties") + sp.getMethod("get", String::class.java).invoke(null, "debug.vela.lowram") as? String + }.getOrNull() + return when (v?.lowercase()) { + "true", "1" -> true + "false", "0" -> false + else -> null + } + } + + /** Register [listener]; returns a handle whose `close()` unregisters. Safe to call any time. */ + fun register(listener: Listener): AutoCloseable { + listeners.add(listener) + return AutoCloseable { listeners.remove(listener) } + } + + /** + * Fan a trim out to every registered holder. Each listener is isolated: one throwing must not + * stop the rest from releasing, since under real pressure we want every byte we can get. + */ + fun dispatch(level: Int) { + Timber.i("MemoryPressure dispatch level=%d listeners=%d", level, listeners.size) + for (l in listeners) { + runCatching { l.release(level) } + .onFailure { Timber.w(it, "MemoryPressure listener failed") } + } + } + + /** + * The app is backgrounded or the OS is genuinely short of memory, so caches that only speed + * things up should go. Everything at or above this level is a "drop it" signal. + */ + fun isSevere(level: Int): Boolean = + level >= ComponentCallbacks2.TRIM_MEMORY_BACKGROUND || + level == ComponentCallbacks2.TRIM_MEMORY_RUNNING_CRITICAL || + level == ComponentCallbacks2.TRIM_MEMORY_RUNNING_LOW + + /** Only the harshest levels, where we drop things that cost real time to rebuild. */ + fun isCritical(level: Int): Boolean = + level >= ComponentCallbacks2.TRIM_MEMORY_COMPLETE || + level == ComponentCallbacks2.TRIM_MEMORY_RUNNING_CRITICAL +} diff --git a/app/src/main/java/app/vela/ui/map/MapViewModel.kt b/app/src/main/java/app/vela/ui/map/MapViewModel.kt index 004d4f4a..134eabf9 100644 --- a/app/src/main/java/app/vela/ui/map/MapViewModel.kt +++ b/app/src/main/java/app/vela/ui/map/MapViewModel.kt @@ -1180,8 +1180,14 @@ class MapViewModel @Inject constructor( // popular times AND the photo gallery land faster when the user taps a result // (both idempotent; the photo warm primes the renderer + HTTP/2 sockets + cache // so the first place page skips the cold start). - viewModelScope.launch { runCatching { webPopularTimes.prewarm() } } - runCatching { webPhotos.warm() } + // Skipped on low-RAM devices: each warm spins up a Chromium renderer SPECULATIVELY, on the + // guess that a search predicts a place tap. When memory is the scarce resource that trade is + // backwards - the user pays two renderers on every search whether or not they open anything + // (issue #83). Those phones build the WebView on first real use instead. + if (!app.vela.ui.MemoryPressure.lowRam) { + viewModelScope.launch { runCatching { webPopularTimes.prewarm() } } + runCatching { webPhotos.warm() } + } searchJob?.cancel() searchJob = viewModelScope.launch { // A fresh typed search leaves any along-route browse: picks open places normally again. @@ -4360,6 +4366,10 @@ class MapViewModel @Inject constructor( } fun deleteAsrModel() { + // Free the loaded model BEFORE removing its files. Deleting the directory alone left the + // native recognizer resident for the rest of the process (~267 MB measured, issue #83), so + // "Remove" reclaimed disk but no memory at all. + whisperRecognizer.release() app.vela.voice.AsrModel.dir(appContext).deleteRecursively() _state.update { it.copy(asrInstalled = false) } } diff --git a/app/src/main/java/app/vela/ui/map/VelaMapView.kt b/app/src/main/java/app/vela/ui/map/VelaMapView.kt index f7ed559d..e93dabe0 100644 --- a/app/src/main/java/app/vela/ui/map/VelaMapView.kt +++ b/app/src/main/java/app/vela/ui/map/VelaMapView.kt @@ -1084,7 +1084,17 @@ fun VelaMapView( } } lifecycleOwner.lifecycle.addObserver(observer) + // MapLibre keeps its tile, glyph and sprite caches in NATIVE memory, and the only way to ask + // it to shrink them is onLowMemory(). Nothing called it before (issue #83), so the map held + // its full cache through every trim the OS sent. Registered with the map's own lifecycle so + // the listener can never outlive the MapView it points at. + val trim = app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isSevere(level)) { + runCatching { mapView.onLowMemory() } + } + } onDispose { + trim.close() lifecycleOwner.lifecycle.removeObserver(observer) mapView.onPause() mapView.onStop() diff --git a/app/src/main/java/app/vela/voice/PiperSynth.kt b/app/src/main/java/app/vela/voice/PiperSynth.kt index 665bc37a..99f30f01 100644 --- a/app/src/main/java/app/vela/voice/PiperSynth.kt +++ b/app/src/main/java/app/vela/voice/PiperSynth.kt @@ -38,6 +38,19 @@ class PiperSynth @Inject constructor( @Volatile private var loadFailed = false @Volatile private var generation = 0 + init { + // The Piper VITS model is the app's second-largest native holding after the ASR model. + // CRITICAL only, deliberately narrower than the recognizer's severe trigger: dropping the + // synth costs a reload on the next prompt, and a prompt arriving late during navigation is + // a missed turn. TRIM_MEMORY_COMPLETE only reaches background processes, and + // RUNNING_CRITICAL means the device is about to start killing things regardless. + // release() posts to the piper-tts worker, so it is already serialized against an + // in-flight synthesis and cannot free the model out from under one (issue #83). + app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isCritical(level)) release() + } + } + /** Which voice id `tts` currently holds - lets [ensureLoaded] detect a voice switch and rebuild. */ @Volatile private var loadedVoiceId: String? = null diff --git a/app/src/main/java/app/vela/voice/WhisperRecognizer.kt b/app/src/main/java/app/vela/voice/WhisperRecognizer.kt index a7654d1e..1edf95f1 100644 --- a/app/src/main/java/app/vela/voice/WhisperRecognizer.kt +++ b/app/src/main/java/app/vela/voice/WhisperRecognizer.kt @@ -34,9 +34,11 @@ import kotlin.math.sqrt * the end of speech, and returns the transcript. Nothing leaves the phone and no third-party voice * app is needed (that's tier-2 - the RECOGNIZE_SPEECH intent handoff in MapScreen). * - * The Whisper recognizer loads lazily and is kept for the process lifetime (~1 s to load); the VAD is - * created per listen (it's tiny and holds streaming state). R8 must keep `com.k2fsa.sherpa.onnx.**` - * (JNI resolves classes by name) - already in `consumer-rules`/`proguard` for Piper. + * The Whisper recognizer loads lazily (~1 s) and is NO LONGER kept for the process lifetime: it costs + * ~267 MB PSS, so it is dropped after [REAP_IDLE_MS] of quiet and on any severe memory trim, and + * rebuilt on the next use (issue #83). The VAD is created per listen (it's tiny and holds streaming + * state). R8 must keep `com.k2fsa.sherpa.onnx.**` (JNI resolves classes by name) - already in + * `consumer-rules`/`proguard` for Piper. */ @Singleton class WhisperRecognizer @Inject constructor( @@ -46,6 +48,67 @@ class WhisperRecognizer @Inject constructor( @Volatile private var recognizer: OfflineRecognizer? = null @Volatile private var loadedLang: String? = null + /** Non-zero while a [listen] is inside the native recognizer. [release] refuses to free the + * model while this is set: `OfflineRecognizer.release()` frees C++ memory that an in-flight + * decode is still reading, which is a use-after-free that takes the process down rather than + * throwing. A trim arriving mid-utterance simply keeps the model until the utterance ends. */ + private val inFlight = java.util.concurrent.atomic.AtomicInteger(0) + + /** Idle-reap timer, same idea as the web fetchers' `REAP_IDLE_MS` (issue #182). One daemon + * thread, shared, created lazily so a device that never loads the model never starts it. */ + private val reaper by lazy { + java.util.concurrent.Executors.newSingleThreadScheduledExecutor { r -> + Thread(r, "asr-reaper").apply { isDaemon = true } + } + } + @Volatile private var reapTask: java.util.concurrent.ScheduledFuture<*>? = null + + init { + // Measured on an M5 (2.9 GB, standardDebug, issue #83): the loaded Whisper tiny int8 model + // costs ~267 MB PSS - ~101 MB of weights in scudo:secondary plus ~146 MB of onnxruntime + // arena in scudo:primary. That was resident for the whole process with no way to reclaim it, + // and it survived deleteAsrModel(). It is by far the largest single reclaimable allocation + // in the app, so it releases on any severe trim and reloads (~1 s) on the next listen. + app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isSevere(level)) release() + } + } + + /** + * Drop the model after a quiet period, on EVERY device, not just low-RAM ones. + * + * Warming at startup buys an instant first mic tap (a user asked for it, 2026-07-10) but the app + * was then holding ~267 MB for the whole session on the CHANCE of a tap that many users never + * make. Reaping after idle keeps the instant first tap and stops the model outliving the user's + * interest in it; a later tap pays the same ~1 s load the very first one used to. Every load and + * every listen re-arms the timer, so an active dictation session never reaps mid-use. + */ + private fun armIdleReap() { + reapTask?.cancel(false) + reapTask = runCatching { + reaper.schedule({ release() }, REAP_IDLE_MS, java.util.concurrent.TimeUnit.MILLISECONDS) + }.getOrNull() + } + + /** + * Free the native recognizer. Safe to call any time: no-op when nothing is loaded, and declines + * while a listen is in flight (see [inFlight]). The next [listen]/[warmUp] rebuilds it. + */ + fun release() { + if (inFlight.get() > 0) { + Timber.tag(TAG).i("release skipped, listen in flight") + return + } + synchronized(loadLock) { + val r = recognizer ?: return + recognizer = null + loadedLang = null + runCatching { r.release() } + .onFailure { Timber.tag(TAG).w(it, "recognizer release failed") } + Timber.tag(TAG).i("recognizer released") + } + } + private val audioManager by lazy { context.getSystemService(Context.AUDIO_SERVICE) as? AudioManager } @Volatile private var focusRequest: AudioFocusRequest? = null @@ -85,6 +148,10 @@ class WhisperRecognizer @Inject constructor( const val SAMPLE_RATE = 16000 const val VAD_WINDOW = 512 // Silero v4/v5 window at 16 kHz const val MAX_SECONDS = 15 // hard cap on one utterance + // Drop the loaded model after this quiet period (issue #83). Matches the web + // fetchers' REAP_IDLE_MS: long enough that a dictation session never reaps between + // utterances, short enough that a session-long 267 MB hold cannot happen. + const val REAP_IDLE_MS = 120_000L } fun isInstalled(): Boolean = @@ -123,6 +190,14 @@ class WhisperRecognizer @Inject constructor( * lazily on the next listen (rare enough not to chase). */ fun warmUp() { if (!AsrModel.isInstalled(context)) return + // On a low-RAM device the warm-up is a bad trade: it spends ~267 MB (measured, issue #83) at + // EVERY launch to save ~1 s on a mic tap the user may never make, and refreshAsr() calls this + // from VM init plus two LaunchedEffects. Those phones load on first listen instead. Roomier + // devices keep the instant-mic behaviour they have always had. + if (app.vela.ui.MemoryPressure.lowRam) { + Timber.tag(TAG).i("skipping ASR warm-up on a low-RAM device, will load on first listen") + return + } Thread({ runCatching { ensureRecognizer() } }, "asr-warmup").start() } @@ -184,6 +259,7 @@ class WhisperRecognizer @Inject constructor( prefs.edit().putBoolean(KEY_LOAD_INFLIGHT, false).apply() recognizer = r loadedLang = lang + if (r != null) armIdleReap() // start the quiet-period countdown from the load return r } } @@ -195,11 +271,30 @@ class WhisperRecognizer @Inject constructor( * (the user tapped done/close). Runs off the main thread; safe to cancel via coroutine too. */ /** Listen, transcribe, and say WHY when it does not work - see [VoiceResult]. Every failure exit - * logs under `VELAASR` so a tester's logcat names the cause without another round-trip. */ + * logs under `VELAASR` so a tester's logcat names the cause without another round-trip. + * + * Thin wrapper over [listenInner] that marks the recognizer busy, so a memory trim arriving + * mid-utterance cannot free the native model out from under the decode (see [inFlight]). The + * inner function keeps its many early returns; this keeps the guard exception-safe. */ suspend fun listen( onLevel: (Float) -> Unit, onListening: () -> Unit, cancelled: () -> Boolean, + ): VoiceResult { + inFlight.incrementAndGet() + reapTask?.cancel(false) // never reap mid-utterance + try { + return listenInner(onLevel, onListening, cancelled) + } finally { + inFlight.decrementAndGet() + armIdleReap() // restart the quiet period from the END of this utterance + } + } + + private suspend fun listenInner( + onLevel: (Float) -> Unit, + onListening: () -> Unit, + cancelled: () -> Boolean, ): VoiceResult = withContext(Dispatchers.Default) { fun fail(reason: VoiceResult.Reason, detail: String? = null): VoiceResult.Failed { Timber.tag(TAG).e("listen failed: $reason${detail?.let { " ($it)" } ?: ""}") diff --git a/app/src/main/java/app/vela/web/WebDirectionsFetcher.kt b/app/src/main/java/app/vela/web/WebDirectionsFetcher.kt index 98be24ff..fbf2b053 100644 --- a/app/src/main/java/app/vela/web/WebDirectionsFetcher.kt +++ b/app/src/main/java/app/vela/web/WebDirectionsFetcher.kt @@ -64,14 +64,26 @@ class WebDirectionsFetcher @Inject constructor( * real memory. The next fetch after a reap just re-creates it. */ private fun scheduleReap() { reap?.let(main::removeCallbacks) - val r = Runnable { - webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } - webView = null - } + val r = Runnable { reapNow() } reap = r main.postDelayed(r, REAP_IDLE_MS) } + /** Destroy the WebView immediately. Must run on the main thread (WebView requirement). */ + private fun reapNow() { + webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } + webView = null + } + + init { + // Under real memory pressure the 120 s idle timer is far too slow - the OS is asking for + // memory NOW and a Chromium renderer is one of the largest things we hold (issue #83). + // Reap on the main thread, since WebView.destroy() requires it. + app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { cancelReap(); reapNow() } + } + } + private fun cancelReap() { reap?.let(main::removeCallbacks) reap = null diff --git a/app/src/main/java/app/vela/web/WebPhotoFetcher.kt b/app/src/main/java/app/vela/web/WebPhotoFetcher.kt index 3e2b5669..7985f467 100644 --- a/app/src/main/java/app/vela/web/WebPhotoFetcher.kt +++ b/app/src/main/java/app/vela/web/WebPhotoFetcher.kt @@ -54,6 +54,25 @@ class WebPhotoFetcher @Inject constructor( @Volatile private var webView: WebView? = null @Volatile private var warmed = false + init { + // Unlike the other four web fetchers this one has NO idle reaper, so a warmed gallery + // renderer was pinned for the whole session with only renderer-death to clear it. A + // Chromium renderer is one of the largest things the app holds, and this fetcher is one of + // the two warmed speculatively on every search, so it releases under pressure (issue #83). + app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { reapNow() } + } + } + + /** Destroy the WebView immediately. Main thread only (WebView requirement). The next + * [warm]/fetch rebuilds it via `ensureWebView`, exactly as after a renderer death. */ + private fun reapNow() { + val wv = webView ?: return + webView = null + warmed = false + runCatching { wv.loadUrl("about:blank"); wv.destroy() } + } + // featureId → its scraped gallery. Re-tapping a place (or bouncing back from directions) then // shows photos INSTANTLY instead of re-running the ~20 s scrape. Access-order LRU, small cap. private val cache = object : LinkedHashMap>(16, 0.75f, true) { diff --git a/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt b/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt index dff125a7..90e77ae4 100644 --- a/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt +++ b/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt @@ -59,15 +59,27 @@ class WebPopularTimesFetcher @Inject constructor( * reap re-warms (google.com -> maps), a one-off few-second cost after minutes idle. */ private fun scheduleReap() { reap?.let(main::removeCallbacks) - val r = Runnable { - webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } - webView = null - warm = null // ensureWarm re-runs the warm sequence on the next fetch - } + val r = Runnable { reapNow() } reap = r main.postDelayed(r, REAP_IDLE_MS) } + /** Destroy the WebView immediately. Must run on the main thread (WebView requirement). */ + private fun reapNow() { + webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } + webView = null + warm = null // ensureWarm re-runs the warm sequence on the next fetch + } + + init { + // Under real memory pressure the 120 s idle timer is far too slow - the OS is asking for + // memory NOW and a Chromium renderer is one of the largest things we hold (issue #83). + // Reap on the main thread, since WebView.destroy() requires it. + app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { cancelReap(); reapNow() } + } + } + private fun cancelReap() { reap?.let(main::removeCallbacks) reap = null diff --git a/app/src/main/java/app/vela/web/WebReviewsFetcher.kt b/app/src/main/java/app/vela/web/WebReviewsFetcher.kt index d25bf7cc..051bb490 100644 --- a/app/src/main/java/app/vela/web/WebReviewsFetcher.kt +++ b/app/src/main/java/app/vela/web/WebReviewsFetcher.kt @@ -59,14 +59,26 @@ class WebReviewsFetcher @Inject constructor( * real memory. The next fetch after a reap just re-creates it. */ private fun scheduleReap() { reap?.let(main::removeCallbacks) - val r = Runnable { - webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } - webView = null - } + val r = Runnable { reapNow() } reap = r main.postDelayed(r, REAP_IDLE_MS) } + /** Destroy the WebView immediately. Must run on the main thread (WebView requirement). */ + private fun reapNow() { + webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } + webView = null + } + + init { + // Under real memory pressure the 120 s idle timer is far too slow - the OS is asking for + // memory NOW and a Chromium renderer is one of the largest things we hold (issue #83). + // Reap on the main thread, since WebView.destroy() requires it. + app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { cancelReap(); reapNow() } + } + } + private fun cancelReap() { reap?.let(main::removeCallbacks) reap = null diff --git a/app/src/main/java/app/vela/web/WebStopDeparturesFetcher.kt b/app/src/main/java/app/vela/web/WebStopDeparturesFetcher.kt index dc3a4edd..32d1dff4 100644 --- a/app/src/main/java/app/vela/web/WebStopDeparturesFetcher.kt +++ b/app/src/main/java/app/vela/web/WebStopDeparturesFetcher.kt @@ -54,14 +54,26 @@ class WebStopDeparturesFetcher @Inject constructor( * just re-creates it - a one-off warm-up, only after minutes of not using the feature. */ private fun scheduleReap() { reap?.let(main::removeCallbacks) - val r = Runnable { - webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } - webView = null - } + val r = Runnable { reapNow() } reap = r main.postDelayed(r, REAP_IDLE_MS) } + /** Destroy the WebView immediately. Must run on the main thread (WebView requirement). */ + private fun reapNow() { + webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } + webView = null + } + + init { + // Under real memory pressure the 120 s idle timer is far too slow - the OS is asking for + // memory NOW and a Chromium renderer is one of the largest things we hold (issue #83). + // Reap on the main thread, since WebView.destroy() requires it. + app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { cancelReap(); reapNow() } + } + } + private fun cancelReap() { reap?.let(main::removeCallbacks) reap = null diff --git a/core/src/main/java/app/vela/core/data/LowRamMode.kt b/core/src/main/java/app/vela/core/data/LowRamMode.kt new file mode 100644 index 00000000..56d7272d --- /dev/null +++ b/core/src/main/java/app/vela/core/data/LowRamMode.kt @@ -0,0 +1,17 @@ +package app.vela.core.data + +/** + * Whether this device is memory-constrained, exposed as a `:core`-visible flag. + * + * Same shape and reason as [CategoryFilter.enabled]: the detection lives in `:app` + * (`app.vela.ui.MemoryPressure`, which needs `ActivityManager`), but the behaviour it gates has to + * act down at the data-source seam. `:core` stays UI-agnostic and never reads an app holder, so the + * app pushes the value in at startup instead. + * + * Off by default, which keeps every roomier device byte-identical to previous behaviour. + */ +object LowRamMode { + + /** Set once from `VelaApp.onCreate` after `MemoryPressure.init`. */ + @Volatile var enabled: Boolean = false +} diff --git a/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt b/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt index facacd4d..6f121bb9 100644 --- a/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt +++ b/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt @@ -6,6 +6,7 @@ import app.vela.core.config.JsTransforms import app.vela.core.diag.DiagLog import app.vela.core.data.CalibrationNeededException import app.vela.core.data.CategoryFilter +import app.vela.core.data.LowRamMode import app.vela.core.data.MapDataSource import app.vela.core.data.RouteEngine import app.vela.core.data.RouteGeometry @@ -142,7 +143,7 @@ class GoogleMapsDataSource @Inject constructor( // FAN OUT across category terms + merge: one "places" query is biased to prominent food/ // shops, so it misses whole tiers (a strip mall's plumber, nail salon, IT shop). A handful // of category queries roughly DOUBLES local coverage (live: 22→52 unique within 600 m). - val terms = listOf( + val allTerms = listOf( "places", "restaurants", "coffee", "stores", "shopping", "services", "beauty salon", "fast food", // High-traffic everyday categories the food/shop-biased set above under-returns, so the map // shows a Google-like MIX (a gas station, a gym, a grocer) rather than mostly restaurants. @@ -154,12 +155,36 @@ class GoogleMapsDataSource @Inject constructor( // their prominence low, so they surface in quiet/residential views without crowding businesses. "school", "park", ) + // LOW-RAM: the fan-out is the app's single largest allocation burst. Each term buffers a + // full response String, a stripped copy, and a JsonElement DOM (GoogleResponse.parse), and + // 15 of those run 4-at-a-time per pan - the ~180 MB/12 s churn in AGENTS.md's memory rule. + // Constrained devices fetch an 8-term subset instead of 15. + // + // The subset is NOT just the first N. "school" and "park" are retained DELIBERATELY: while + // the ambient layer is active at z14+ the basemap's OSM poi layers are filter-hidden, so + // those categories have NO second source - dropping them makes parks and schools vanish + // from the map entirely, which is the exact bug the civic/green terms were added to fix + // (see the comment on allTerms above). Device-verified by A/B screenshot: a first attempt + // at a 6-term subset lost every park and school pin on the low-RAM frame. + // + // What goes instead are the terms whose places still surface via "places"/"stores" or whose + // absence degrades gracefully: shopping, services, beauty salon, fast food, gym, bar, + // pharmacy. Fewer ambient POIs is a visible trade, and the right one on a phone that + // otherwise OOMs (issue #83). Roomier devices are unaffected. + val terms = if (LowRamMode.enabled) { + listOf("places", "restaurants", "coffee", "stores", "grocery store", "gas station", "school", "park") + } else { + allTerms + } suspend fun fetchTerm(term: String): List = ambientFanout.withPermit { runCatching { val pb = SearchPb.build(term, center, cal.searchPb) .replaceFirst(Regex("!1d[0-9.]+"), "!1d${spanMeters.toInt()}") .replaceFirst(Regex("!4f[0-9.]+"), "!4f${String.format(java.util.Locale.US, "%.1f", zoom)}") - .replaceFirst(Regex("!7i\\d+"), "!7i60") // deep pool per term, so zooming in can go down the rank + // Deep pool per term, so zooming in can go down the rank. Halved on low-RAM: the + // pool size drives the RESPONSE BODY size, and the body is what gets buffered + // and DOM-parsed per term (issue #83). + .replaceFirst(Regex("!7i\\d+"), if (LowRamMode.enabled) "!7i30" else "!7i60") val url = "${cal.searchEndpoint}&q=${term.enc()}&pb=${pb.enc()}".localized() SearchParser.parse(term, GoogleResponse.parse(get(url)), center, cal.paths).places }.getOrDefault(emptyList()) diff --git a/docs/FEATURES.md b/docs/FEATURES.md index ec30bbf9..d4fe2235 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -284,6 +284,8 @@ Status legend: [x] done · [~] partial / in progress · [ ] planned - [x] CI builds, tests, signs and publishes a normal release `v0.0.` with debug and release APKs; no prerelease channel, tracked by Obtainium and the updater with zero config. - [x] **Opt-in diagnostics/debug export** (Settings → Diagnostics, off by default) - a local-only event log exportable to JSON via the share sheet, never auto-uploaded, wiped when turned off, in-memory only. - [x] **Crash/ANR/jank capture, all local** - an uncaught-exception handler persists stack traces + breadcrumbs to disk for export, ApplicationExitInfo harvests ANR/native/low-memory kills, a debug ANR watchdog and StrictMode flag stalls and main-thread I/O; captured even with diagnostics off, never auto-sent. +- [x] **Memory use cut across the board** (issue #83) - the on-device speech model costs ~267 MB while loaded and used to stay resident for the entire session; it is now dropped after two minutes unused and rebuilt on next use, which reclaims ~101 MB on ANY phone at no visible cost. Every large or native allocation also releases when the system reports memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and gave back nothing when asked. Measured on a 2.9 GB test device: idle memory down 29%, post-trim down 38%, native heap down 57%. +- [x] **Low-RAM device adaptation** (issue #83) - additionally, the app detects a memory-constrained phone (`isLowRamDevice` or a heap class of 128 MB or less) and adapts: the on-device speech model is not preloaded at startup, the image cache is capped at 16 MB instead of 48 MB, WebView renderers are not warmed speculatively, and the ambient POI fan-out drops from 15 category terms to 8 with a smaller result pool. Independently, every large or native allocation now releases on OS memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and returned nothing when the OS asked. Measured on a 2.9 GB test device: peak memory down 30%, post-trim memory down 38%, native heap down 57%, cold start ~480 ms faster. x86 and x86_64 native libraries are no longer shipped (no target phone can run them), cutting 23 MB of dead weight from every install. - [x] **Trip recording + replay** (Settings → "Save my trips", off by default, separate opt-in) - records each drive's GPS trace to a local file replayable on the map at 3x through the real nav pipeline; saved on arrival, listed with Replay/Share/Delete, Share exporting the raw CSV; replay auto-routes to the destination. - [x] **Simulate driving (demo mode)** (Settings → Navigation, off by default) - Start drives any planned route as a synthetic GPS trace through the live-nav loop so nav runs anywhere for demos and screenshots; End stops it; turn off to navigate for real. - [x] **Simulate my location (demo mode)** (off by default) - Vela pretends you're at the map centre so the dot, directions origin and recenter read from there; turn off for real GPS. From 7345474a84bb75fa85fb35de88275d9ff969f3c4 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Mon, 20 Jul 2026 22:29:50 -0400 Subject: [PATCH 02/21] Hand freed memory back to the OS, not just to the allocator PR #85 gave every big holder a release() and fanned OS trims out to them, but a Kotlin release only returns pages to scudo. They sit on its free lists, where RSS/PSS still count them and lmkd still sees a fat process. Measured on the M5 before this change: a full TRIM_MEMORY_COMPLETE with all 8 listeners firing moved scudo:primary 56,578 to 54,978 KB while mallinfo reported a 442 MB arena holding just 46 MB live. mallopt() is the only way to hand that gap on and it is reachable only from C, so this adds the app's first native module: app/src/main/cpp/velamem.cpp, three lines calling libc, 4 KB for arm64 and 2.7 KB for armeabi-v7a. Built for those two ABIs only, matching the x86 drop in #85. MemoryPressure.dispatch schedules it 750 ms after a trim, off the main thread. The delay is load-bearing: the WebView reapers post destroy() to the main looper and VelaApp clears Coil after dispatch returns, so an inline purge would run before the memory it is meant to reclaim had been freed. The purge fires from TRIM_MEMORY_RUNNING_LOW (10) up, deliberately wider than isSevere (40). Measured: pressing HOME delivers only TRIM_MEMORY_UI_HIDDEN (20), never BACKGROUND (40), so gating on isSevere would skip the single most common moment we are handed, the one where the app is off-screen and nothing can jank. Verified by A/B on ONE binary, since two builds also differ in background settling and idle PSS swings +-60 MB run to run. debug.vela.nopurge suppresses the purge at runtime, making the delta paired within a run; both arms were checked in logcat to confirm the gate actually gates. Releasing the ASR model, 3 alternating pairs: with the purge suppressed scudo:primary moved 60/32/28 KB in the 7 s after the trim, which is nothing, and with it on 3704/3008/2792 KB. No overlap. A second 8-pair A/B over map and POI churn agreed: 3345 to 6931 KB mean reclaimed, Mann-Whitney U=7 at n=8/8, p<0.05. It is worth a consistent ~3 MB, not tens, and the commit says so rather than claiming the onnxruntime arena. The ASR model's ~111 MB lives in scudo:secondary, which is mmap-backed and comes back on free() with no purge needed (111 MB to 7 MB in BOTH arms). Only scudo:primary needs asking. M_PURGE_ALL is API 34+, so on the Android 13 dev phone it returns 0 and the code falls back to M_PURGE (API 28+). The logged mode= says which actually took, so a device supporting neither is visible instead of silently doing nothing. Also corrects two things #85 left wrong. The MemoryPressure KDoc told the reader to run `setprop debug.vela.lowram ""` for real detection, which AGENTS.md itself says is a shell syntax error; and `false` does not restore detection either, it forces the normal path. Clearing needs an unparseable value, so both the KDoc and AGENTS.md now say `none`, which is verified: the app logs forced=no. AGENTS.md also gains the UI_HIDDEN-vs-BACKGROUND finding, the one-binary A/B rule, and the asr_model_bad trap: a quarantined model makes warmUp() a silent no-op, so scudo:secondary sits at ~11 MB instead of ~111 MB and a memory benchmark measures the model-absent case without saying so. That cost a run here. Verified: :app:detekt and :core:detekt 0 smells, :core:test green, audit_deadcode.sh PASS, standardDebug builds, installs, launches with no UnsatisfiedLinkError, and renders correctly on device (screenshot: map, POI pins including parks, focus ring, soft keys). --- AGENTS.md | 44 +++++++- app/build.gradle.kts | 17 +++ app/src/main/cpp/CMakeLists.txt | 12 ++ app/src/main/cpp/velamem.cpp | 38 +++++++ .../main/java/app/vela/ui/MemoryPressure.kt | 104 +++++++++++++++++- docs/FEATURES.md | 2 +- 6 files changed, 208 insertions(+), 9 deletions(-) create mode 100644 app/src/main/cpp/CMakeLists.txt create mode 100644 app/src/main/cpp/velamem.cpp diff --git a/AGENTS.md b/AGENTS.md index af526d06..fadad6df 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -649,8 +649,10 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): as `CategoryFilter.enabled`, because `:core` must never read an `:app` holder. - **Verify the low-RAM path or it ships unverified.** Every dev phone we own reports `lowRam=false heapClassMb=256`, so those branches are dead code locally. Debug builds honour - `adb shell setprop debug.vela.lowram true` (then relaunch); `setprop debug.vela.lowram false` - restores real detection. NB `setprop ""` is a syntax error, not a reset. + `adb shell setprop debug.vela.lowram true` (then relaunch). NB `false` FORCES the normal path, + it does NOT clear the override - clearing needs an unparseable value, so use + `setprop debug.vela.lowram none`. The two only look equivalent because every dev phone we own + detects as normal anyway. `setprop ""` is a shell syntax error, not a reset. - **Measuring: `am send-trim-memory` REFUSES background levels on a foreground process** ("Unable to set a background trim level on a foreground process"). Press HOME first. A harness that discards that stderr measures NOTHING and reports a clean baseline - this happened here and @@ -677,6 +679,44 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): subset keeps `school` and `park` on purpose: the ambient layer filter-hides the basemap OSM poi layers at z14+, so those two would vanish entirely. A first 6-term subset did exactly that and was caught by an A/B screenshot, not by any test. + - **A Kotlin `release()` does NOT give memory back to the KERNEL, only to scudo.** Freeing a model + or a WebView returns its pages to the allocator's free lists, where RSS/PSS still count them and + lmkd still sees a fat process. `mallopt()` is the only way to hand them on and it is reachable + only from C, which is why `app/src/main/cpp/velamem.cpp` exists - the app's ONLY native module, + ~4 KB per ABI, built for `arm64-v8a` + `armeabi-v7a` only. `MemoryPressure.dispatch` schedules + it 750 ms after every trim (the delay matters: the WebView reapers post `destroy()` to the main + looper and `VelaApp` clears Coil after `dispatch` returns, so an inline purge would run before + the memory it is meant to reclaim was actually freed). + - Measured on the M5, 3 alternating A/B pairs, ASR-model-release scenario: with the purge + suppressed `scudo:primary` moved 60/32/28 KB in the 7 s after a severe trim, i.e. NOTHING; + with it on, 3704/3008/2792 KB. No overlap. A second 8-pair A/B over map/POI churn showed the + same shape, 3345 KB -> 6931 KB mean reclaimed (Mann-Whitney U=7, n=8/8, p<0.05). + - Do NOT expect this to reclaim the onnxruntime arena. It is worth a consistent ~3 MB, not tens. + The ASR model's ~111 MB lives in `scudo:secondary`, which is mmap-backed and comes back on + `free()` with no purge needed - measured 111 MB -> 7 MB in BOTH arms. The purge only moves + `scudo:primary`, which stays around 75 MB either way. + - `M_PURGE_ALL` is API 34+. On the Android 13 dev phone it returns 0 and the code falls back to + `M_PURGE` (API 28+). The logged `mode=` says which actually took (2, 1, or 0 for neither); + do not assume `M_PURGE_ALL` ran just because `all=true` was passed. + - **Backgrounding delivers `TRIM_MEMORY_UI_HIDDEN` (20), NOT `TRIM_MEMORY_BACKGROUND` (40).** + Measured on the M5: pressing HOME logs `dispatch level=20` and nothing more; 40 arrives only + later, once the process sinks in the LRU list under real pressure. `isSevere` starts at 40, so + it is deliberately FALSE at the single most common moment the app is handed. That is right for + releasing (do not thrash a model on every HOME press) and wrong for purging, which is why the + purge triggers from `TRIM_MEMORY_RUNNING_LOW` (10) up. Anything that should happen "when the + user leaves the app" must key off 20, not 40. + - **A/B a memory change on ONE binary or the arms differ in more than the change.** Two builds + also differ in background settling, and idle PSS swings +-60 MB run to run, so a cross-build + comparison cannot attribute a few MB to anything. `adb shell setprop debug.vela.nopurge true` + suppresses the purge at runtime, which makes the delta a paired within-run measurement. Verify + the gate really gates before trusting either arm: one arm must log `native purge suppressed` + and the other `native purge all=... mode=...`, or the A/B is measuring one thing twice. + - **A quarantined ASR model makes `warmUp()` a silent no-op.** `WhisperRecognizer.isInstalled()` + is `AsrModel.isInstalled() && !asr_model_bad`, and the corrupt-model quarantine (`asr_model_bad` + in `vela_settings`) is only ever lifted by the installer's `clearQuarantine()`. Side-loading the + model files by hand leaves the flag set, so the model never loads, `scudo:secondary` sits at + ~11 MB instead of ~111 MB, and a memory benchmark silently measures the model-absent case. Check + `secondary` is actually ~111 MB before believing any ASR memory number. ## Layout diff --git a/app/build.gradle.kts b/app/build.gradle.kts index 8194c01a..0a4cfdb9 100644 --- a/app/build.gradle.kts +++ b/app/build.gradle.kts @@ -49,6 +49,13 @@ android { testInstrumentationRunner = "androidx.test.runner.AndroidJUnitRunner" vectorDrawables { useSupportLibrary = true } + // libvelamem: the mallopt() purge shim (app/src/main/cpp). Built ONLY for the two ABIs the + // app actually ships - the packaging block below drops x86/x86_64 for every other native + // lib, so building them here would just produce artifacts that get deleted again. + externalNativeBuild { + cmake { abiFilters += listOf("arm64-v8a", "armeabi-v7a") } + } + // MapTiler key injected from the CI secret (-PmaptilerKey); empty for // local builds, in which case the app falls back to the keyless // OpenFreeMap basemap. Never stored in the repo. @@ -193,6 +200,16 @@ android { } } + // The app's only native code: the mallopt() purge shim. Pinned NDK/CMake versions so a + // developer with a different NDK installed gets the same libvelamem.so as CI. + ndkVersion = "27.0.12077973" + externalNativeBuild { + cmake { + path = file("src/main/cpp/CMakeLists.txt") + version = "3.22.1" + } + } + compileOptions { sourceCompatibility = JavaVersion.VERSION_17 targetCompatibility = JavaVersion.VERSION_17 diff --git a/app/src/main/cpp/CMakeLists.txt b/app/src/main/cpp/CMakeLists.txt new file mode 100644 index 00000000..5f9808c4 --- /dev/null +++ b/app/src/main/cpp/CMakeLists.txt @@ -0,0 +1,12 @@ +# The app's only native module: a mallopt() purge shim (see velamem.cpp for why). +# Deliberately tiny and STL-free, so it adds a few KB per ABI rather than a runtime. +cmake_minimum_required(VERSION 3.22.1) +project(velamem LANGUAGES CXX) + +add_library(velamem SHARED velamem.cpp) + +# -Os and no exceptions/RTTI: this is three lines of code calling libc, none of which needs them. +target_compile_options(velamem PRIVATE -Os -fno-exceptions -fno-rtti -fvisibility=hidden) + +# Only libc is needed; mallopt lives there. No log lib, the Kotlin side does the logging. +target_link_libraries(velamem) diff --git a/app/src/main/cpp/velamem.cpp b/app/src/main/cpp/velamem.cpp new file mode 100644 index 00000000..6db7428c --- /dev/null +++ b/app/src/main/cpp/velamem.cpp @@ -0,0 +1,38 @@ +// Native-allocator purge. The ONLY reason this module exists. +// +// Issue #83 gave every big holder a release() and fanned OS trims out to them (MemoryPressure), but +// a Kotlin release only hands pages back to the ALLOCATOR, not to the kernel. Scudo keeps them on +// its free lists, so RSS/PSS barely moves and the OOM killer still sees a fat process. +// +// Measured on the M5 (2.9 GB, Android 13, app.vela.debug, PR #85 build, all 8 listeners firing): +// a full TRIM_MEMORY_COMPLETE moved scudo:primary only 56,578 -> 54,978 KB while mallinfo reported +// a 442 MB arena holding just 46 MB live. That gap is what mallopt() reclaims and nothing on the +// Java side can touch. +// +// bionic exposes exactly one lever for it, and only through libc: +// M_PURGE (API 28+) release free memory in the calling thread's arena +// M_PURGE_ALL (API 34+) walk every arena; documented as able to take 2x+ a plain M_PURGE +// +// The values are ABI-stable, so they are spelled out rather than taken from , which +// keeps the build independent of NDK header vintage. An unsupported option makes mallopt() return +// 0, so calling M_PURGE_ALL on an API 33 device is a harmless no-op that falls through to M_PURGE. + +#include +#include + +#ifndef M_PURGE +#define M_PURGE (-101) +#endif +#ifndef M_PURGE_ALL +#define M_PURGE_ALL (-104) +#endif + +// Returns which lever actually took, so the Kotlin side can log it and a device that supports +// neither is visible in logcat instead of silently doing nothing: +// 2 = M_PURGE_ALL, 1 = M_PURGE, 0 = neither supported +extern "C" JNIEXPORT jint JNICALL +Java_app_vela_ui_MemoryPressure_nativePurge(JNIEnv*, jobject, jboolean all) { + if (all && mallopt(M_PURGE_ALL, 0) != 0) return 2; + if (mallopt(M_PURGE, 0) != 0) return 1; + return 0; +} diff --git a/app/src/main/java/app/vela/ui/MemoryPressure.kt b/app/src/main/java/app/vela/ui/MemoryPressure.kt index 30e8a652..98ffd883 100644 --- a/app/src/main/java/app/vela/ui/MemoryPressure.kt +++ b/app/src/main/java/app/vela/ui/MemoryPressure.kt @@ -55,19 +55,34 @@ object MemoryPressure { /** * Debug-only override so the low-RAM path can be exercised on a normal dev phone: * - * adb shell setprop debug.vela.lowram true # then relaunch the app - * adb shell setprop debug.vela.lowram "" # back to real detection + * adb shell setprop debug.vela.lowram true # force the low-RAM path, then relaunch + * adb shell setprop debug.vela.lowram false # force the normal path, then relaunch + * adb shell setprop debug.vela.lowram none # clear the override, back to real detection + * + * All three need a relaunch; this is read once from [init]. Note that `false` FORCES the normal + * path rather than clearing the override - the two only look alike because every dev phone we + * own detects as normal anyway. Clearing needs an unparseable value ([debugFlag] returns null + * for anything that is not true/1/false/0), hence `none`; `setprop ""` is a shell syntax + * error, not a reset. * * Without this the low-RAM branches are dead code on every device we actually own (the M5 dev - * phone reports heapClassMb=256, lowRam=false), which means they would ship unverified. Returns - * null when unset or on a non-debug build, so release behaviour is untouched. + * phone reports heapClassMb=256, lowRam=false), which means they would ship unverified. + */ + private fun forcedLowRam(): Boolean? = debugFlag("debug.vela.lowram") + + /** + * Read a debug-only tri-state system property. Returns null when unset, unparseable, or on a + * non-debug build, so release behaviour is never affected by one of these. + * + * NB clearing one is `setprop false`, NOT `setprop ""` - an empty value is a + * syntax error at the shell, not a reset. */ - private fun forcedLowRam(): Boolean? { + private fun debugFlag(name: String): Boolean? { if (!app.vela.BuildConfig.DEBUG) return null val v = runCatching { @Suppress("PrivateApi") val sp = Class.forName("android.os.SystemProperties") - sp.getMethod("get", String::class.java).invoke(null, "debug.vela.lowram") as? String + sp.getMethod("get", String::class.java).invoke(null, name) as? String }.getOrNull() return when (v?.lowercase()) { "true", "1" -> true @@ -85,6 +100,9 @@ object MemoryPressure { /** * Fan a trim out to every registered holder. Each listener is isolated: one throwing must not * stop the rest from releasing, since under real pressure we want every byte we can get. + * + * The fan-out alone is only half the job: a listener's `release()` returns pages to SCUDO, not + * to the kernel, so RSS barely moves. [schedulePurge] finishes it. See [nativePurge]. */ fun dispatch(level: Int) { Timber.i("MemoryPressure dispatch level=%d listeners=%d", level, listeners.size) @@ -92,8 +110,82 @@ object MemoryPressure { runCatching { l.release(level) } .onFailure { Timber.w(it, "MemoryPressure listener failed") } } + // Anything except the gentlest level is worth a purge. Deliberately WIDER than isSevere: + // measured on the M5, backgrounding the app delivers only TRIM_MEMORY_UI_HIDDEN (20), never + // TRIM_MEMORY_BACKGROUND (40), so gating the purge on isSevere would skip the single most + // common moment we are handed - the app is off-screen, nothing can jank, and the allocator + // is holding pages nobody will touch again for minutes. + if (level >= ComponentCallbacks2.TRIM_MEMORY_RUNNING_LOW) schedulePurge(isSevere(level)) + } + + // ---------------------------------------------------------------- native allocator purge + + /** + * Scudo hands freed pages back only when asked, and `mallopt` is the only way to ask. Measured + * on the M5 (Android 13, app.vela.debug, all 8 listeners releasing): a full TRIM_MEMORY_COMPLETE + * moved `scudo:primary` just 56,578 -> 54,978 KB while mallinfo showed a 442 MB arena holding + * 46 MB live. Everything in that gap is reclaimable and unreachable from Kotlin. + * + * Returns which lever took: 2 = M_PURGE_ALL, 1 = M_PURGE, 0 = neither (pre-API-28). + */ + private external fun nativePurge(all: Boolean): Int + + /** False when libvelamem is missing (an ABI we do not ship, a stripped install). The purge is + * then skipped rather than taking the process down over an optimization. */ + private val nativeReady: Boolean = + runCatching { System.loadLibrary("velamem") } + .onFailure { Timber.w(it, "libvelamem unavailable, native purge disabled") } + .isSuccess + + /** One daemon thread, created lazily. Never the main thread: a purge walks the allocator's free + * lists behind its global lock, and that is not something to do on the UI thread. */ + private val purgeExec by lazy { + java.util.concurrent.Executors.newSingleThreadScheduledExecutor { r -> + Thread(r, "mem-purge").apply { isDaemon = true } + } + } + + /** Coalesces a burst of trims into one purge. The OS routinely sends several levels in a row. */ + private val purgePending = java.util.concurrent.atomic.AtomicBoolean(false) + + /** + * Purge after [PURGE_DELAY_MS], not immediately: the WebView reapers post their `destroy()` to + * the main looper and `VelaApp` clears Coil right after [dispatch] returns, so an inline purge + * would run BEFORE the memory it is meant to reclaim has actually been freed and reclaim close + * to nothing. + * + * [all] picks the lever. M_PURGE_ALL walks every arena and is documented as able to take over + * twice as long as a plain M_PURGE, so it is spent only on levels [isSevere] already treats as + * "drop it"; a routine UI_HIDDEN gets the cheap one. + */ + private fun schedulePurge(all: Boolean) { + if (!nativeReady) return + // Debug-only kill switch, so the purge's contribution can be A/B measured on ONE binary: + // adb shell setprop debug.vela.nopurge true # then relaunch, this is read per-trim + // Without it the only way to attribute a delta is to compare two different builds, which + // also differ in background settling and cannot be paired inside a single run. + if (debugFlag("debug.vela.nopurge") == true) { + Timber.i("MemoryPressure native purge suppressed by debug.vela.nopurge") + return + } + if (!purgePending.compareAndSet(false, true)) return + val scheduled = runCatching { + purgeExec.schedule({ + purgePending.set(false) + val t0 = android.os.SystemClock.uptimeMillis() + val mode = runCatching { nativePurge(all) }.getOrDefault(0) + Timber.i( + "MemoryPressure native purge all=%b mode=%d took=%dms", + all, mode, android.os.SystemClock.uptimeMillis() - t0, + ) + }, PURGE_DELAY_MS, java.util.concurrent.TimeUnit.MILLISECONDS) + }.getOrNull() + if (scheduled == null) purgePending.set(false) // executor rejected; let the next trim retry } + /** Long enough for the main-looper-posted WebView destroys and the Coil clear to have landed. */ + private const val PURGE_DELAY_MS = 750L + /** * The app is backgrounded or the OS is genuinely short of memory, so caches that only speed * things up should go. Everything at or above this level is a "drop it" signal. diff --git a/docs/FEATURES.md b/docs/FEATURES.md index d4fe2235..be37369f 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -284,7 +284,7 @@ Status legend: [x] done · [~] partial / in progress · [ ] planned - [x] CI builds, tests, signs and publishes a normal release `v0.0.` with debug and release APKs; no prerelease channel, tracked by Obtainium and the updater with zero config. - [x] **Opt-in diagnostics/debug export** (Settings → Diagnostics, off by default) - a local-only event log exportable to JSON via the share sheet, never auto-uploaded, wiped when turned off, in-memory only. - [x] **Crash/ANR/jank capture, all local** - an uncaught-exception handler persists stack traces + breadcrumbs to disk for export, ApplicationExitInfo harvests ANR/native/low-memory kills, a debug ANR watchdog and StrictMode flag stalls and main-thread I/O; captured even with diagnostics off, never auto-sent. -- [x] **Memory use cut across the board** (issue #83) - the on-device speech model costs ~267 MB while loaded and used to stay resident for the entire session; it is now dropped after two minutes unused and rebuilt on next use, which reclaims ~101 MB on ANY phone at no visible cost. Every large or native allocation also releases when the system reports memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and gave back nothing when asked. Measured on a 2.9 GB test device: idle memory down 29%, post-trim down 38%, native heap down 57%. +- [x] **Memory use cut across the board** (issue #83) - the on-device speech model costs ~267 MB while loaded and used to stay resident for the entire session; it is now dropped after two minutes unused and rebuilt on next use, which reclaims ~101 MB on ANY phone at no visible cost. Every large or native allocation also releases when the system reports memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and gave back nothing when asked. Measured on a 2.9 GB test device: idle memory down 29%, post-trim down 38%, native heap down 57%. Freed memory is also handed back to the operating system rather than kept on the allocator's free lists, which a release alone does not do (a measured ~3 MB more returned per memory warning, at no cost to responsiveness). - [x] **Low-RAM device adaptation** (issue #83) - additionally, the app detects a memory-constrained phone (`isLowRamDevice` or a heap class of 128 MB or less) and adapts: the on-device speech model is not preloaded at startup, the image cache is capped at 16 MB instead of 48 MB, WebView renderers are not warmed speculatively, and the ambient POI fan-out drops from 15 category terms to 8 with a smaller result pool. Independently, every large or native allocation now releases on OS memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and returned nothing when the OS asked. Measured on a 2.9 GB test device: peak memory down 30%, post-trim memory down 38%, native heap down 57%, cold start ~480 ms faster. x86 and x86_64 native libraries are no longer shipped (no target phone can run them), cutting 23 MB of dead weight from every install. - [x] **Trip recording + replay** (Settings → "Save my trips", off by default, separate opt-in) - records each drive's GPS trace to a local file replayable on the map at 3x through the real nav pipeline; saved on arrival, listed with Replay/Share/Delete, Share exporting the raw CSV; replay auto-routes to the destination. - [x] **Simulate driving (demo mode)** (Settings → Navigation, off by default) - Start drives any planned route as a synthetic GPS trace through the live-nav loop so nav runs anywhere for demos and screenshots; End stops it; turn off to navigate for real. From 92beb2a0d3e970119fdafc5e549c0667337cbd80 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Tue, 21 Jul 2026 00:08:35 -0400 Subject: [PATCH 03/21] Stop one search from pinning a Chromium renderer for the whole session Measured on the M5 while reviewing #83: the hidden scraper WebViews are the single largest thing Vela costs, and app PSS cannot see most of it. Android runs the WebView renderer OUT OF PROCESS, so after one search sandboxed_process0 sat at 305-347 MB PSS, plus webview_service 21 MB and webview_apk 37 MB, none of it in this app's dumpsys meminfo. Every number in #83 was app-PSS only, so the biggest item in the app was invisible to the whole exercise. The app-side half is GL mtrack, from the offscreen layouts: 26 MB with no WebView alive, 437 MB with the two scraper views up. Neither speculative warm was bounded. WebPhotoFetcher had no idle reaper at all, and WebPopularTimesFetcher.prewarm() created a view and never called scheduleReap() (only fetch() did). MapViewModel warms both on every search, so a single search held the renderer until the process died or a trim arrived. Both now arm a reap, and both matter: fixing only the photo fetcher left the renderer alive at ~200 MB because the popular-times view still held it. The renderer is shared by every WebView in the process, so one un-reaped view keeps it up for all of them. Device-measured, no trim involved anywhere in the run: t+45s GL mtrack 437 MB app PSS 725 MB renderer alive t+90s GL mtrack 65 MB app PSS 337 MB reap logged t+135s GL mtrack 64 MB app PSS 335 MB renderer process gone About 390 MB back to the app plus a ~305 MB renderer process shut down, from memory that used to be held for the rest of the session. Rebuild verified: a later search respawns the renderer and GL mtrack returns to 431 MB, no crash. A speculative warm gets a longer window than a real fetch (WARM_REAP_IDLE_MS 300 s vs REAP_IDLE_MS 120 s). The warm exists so the first place tap skips the cold start, and reaping at 120 s would expire during an ordinary browse and waste it. Bounded, not short, is the point: the bug was session-long. reapNow() now drains pending like rendererGone() already did. Destroying the view kills the injected scraper, so nothing completes those deferreds; a reap landing mid-fetch parked the fetch in deferred.await() for the full 40 s TOTAL_TIMEOUT_MS while holding the fetcher's Mutex, stalling everything queued behind it. An empty result is the documented best-effort failure, a 40 s hang is not. This was already reachable from #83's severe-trim hook. Reap bookkeeping moved onto the main thread in both fetchers (onMain). The reap field is touched from the trim listener on main and from fetch/warm on the caller's dispatcher, and scheduling is a read-modify-write that @Volatile would not make safe; this commit adds another writer, so it fixes the race rather than widening it. The reap is logged because verifying it needs a log, not a process check. WebView.destroy() does not kill the renderer promptly (220 MB still resident 8 s after a destroy, process gone only minutes later) and an OS trim can kill it for unrelated reasons. The first attempt to verify this was unfalsifiable for that reason: the renderer vanished at t+60 s and the logs showed dispatch level=15/40, a real trim rather than the reaper. The runs above assert the reap log AND assert no severe trim fired. AGENTS.md: the out-of-process measurement rule (always total the WebView processes), the both-warms-or-neither finding, the drain-pending rule, and the log-not-process verification rule. Verified: :app:detekt and :core:detekt 0 smells, :core:test green, audit_deadcode.sh PASS, audit_dpad static no new violations (the four reported are pre-existing, in files this commit does not touch), standardDebug builds, installs and runs, screenshot of the map after a reap-and-rebuild cycle. --- AGENTS.md | 32 ++++++++ .../main/java/app/vela/web/WebPhotoFetcher.kt | 74 +++++++++++++++++-- .../app/vela/web/WebPopularTimesFetcher.kt | 30 ++++++-- docs/FEATURES.md | 2 +- 4 files changed, 126 insertions(+), 12 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index fadad6df..61902f49 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -711,6 +711,38 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): suppresses the purge at runtime, which makes the delta a paired within-run measurement. Verify the gate really gates before trusting either arm: one arm must log `native purge suppressed` and the other `native purge all=... mode=...`, or the A/B is measuring one thing twice. + - **The hidden WebViews are the single largest thing Vela costs, and app PSS CANNOT SEE IT.** + Android runs the WebView renderer OUT OF PROCESS. Measured on the M5 after one search: + `sandboxed_process0` at 305-347 MB PSS, plus `webview_service` 21 MB and `webview_apk` 37 MB, + none of it in the app's own `dumpsys meminfo`. Every issue-#83 number was app-PSS only, so the + largest item in the app was invisible to the whole exercise. **When measuring memory here, + always `adb shell ps -A | grep sandboxed_process` and total the WebView processes too.** + - The app-side half is `GL mtrack`, and it is huge: 26 MB with no WebView alive, 396-461 MB once + the two scraper WebViews exist. That is the offscreen layouts (`WV_WIDTH`x`WV_HEIGHT`, e.g. + 1200x3200) allocating graphics buffers charged to OUR process. Chromium logs + `tile memory limits exceeded` at that size. Cutting the offscreen viewport is a real lead. + - **BOTH speculative warms have to be bounded or neither helps.** The renderer is SHARED by every + WebView in the process, so one un-reaped view keeps it alive for all of them. `WebPhotoFetcher` + had no idle reaper at all, and `WebPopularTimesFetcher.prewarm()` created a view and never + called `scheduleReap()` (only `fetch()` did). `MapViewModel` warms both on every search, so one + search pinned the renderer for the session. Device-verified the hard way: reaping only the photo + view left the renderer alive at ~200 MB because the popular-times view still held it. Fixing + both, with no trim involved, took the renderer to zero and app PSS 744 MB -> 351 MB. + - A speculative warm gets a LONGER reap window than a real fetch (`WARM_REAP_IDLE_MS` 300 s vs + `REAP_IDLE_MS` 120 s). The warm exists so the first place tap skips the cold start; reaping it + at 120 s would expire during an ordinary browse and waste the warm entirely. Bounded, not + short, is the goal - the bug was session-long, not "not aggressive enough". + - **A reap must drain `pending`, exactly like `rendererGone` does.** Destroying the view kills the + injected scraper, so nothing will ever complete those deferreds; a reap landing mid-fetch parks + the fetch in `deferred.await()` for the full `TOTAL_TIMEOUT_MS` (40 s) while it HOLDS the + fetcher's `Mutex`, stalling everything queued behind it. An empty result is the documented + best-effort failure; a 40 s hang is not. + - **Verifying a reap needs a LOG, not a process check.** `WebView.destroy()` does not kill the + renderer promptly - measured 220 MB still resident 8 s after a destroy and the process gone only + minutes later - and an OS trim can kill it for unrelated reasons, so "the process went away" does + not mean your timer fired. The first attempt to verify this was unfalsifiable for exactly that + reason: the renderer vanished at t+60 s and the logs showed `dispatch level=15`/`40`, i.e. a real + trim, not the reaper. Log the reap, then assert the log AND assert no severe trim fired. - **A quarantined ASR model makes `warmUp()` a silent no-op.** `WhisperRecognizer.isInstalled()` is `AsrModel.isInstalled() && !asr_model_bad`, and the corrupt-model quarantine (`asr_model_bad` in `vela_settings`) is only ever lifted by the installer's `clearQuarantine()`. Side-loading the diff --git a/app/src/main/java/app/vela/web/WebPhotoFetcher.kt b/app/src/main/java/app/vela/web/WebPhotoFetcher.kt index 7985f467..eaf6a8dd 100644 --- a/app/src/main/java/app/vela/web/WebPhotoFetcher.kt +++ b/app/src/main/java/app/vela/web/WebPhotoFetcher.kt @@ -54,23 +54,73 @@ class WebPhotoFetcher @Inject constructor( @Volatile private var webView: WebView? = null @Volatile private var warmed = false + /** Pending idle-reap callback. Only ever read or written on the main thread - see [onMain]. */ + private var reap: Runnable? = null + init { - // Unlike the other four web fetchers this one has NO idle reaper, so a warmed gallery - // renderer was pinned for the whole session with only renderer-death to clear it. A - // Chromium renderer is one of the largest things the app holds, and this fetcher is one of - // the two warmed speculatively on every search, so it releases under pressure (issue #83). + // A trim is the OS asking for memory NOW, far sooner than any idle timer (issue #83). app.vela.ui.MemoryPressure.register { level -> - if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { reapNow() } + if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { cancelReap(); reapNow() } } } + /** + * Run [block] on the main thread, inline when already there. + * + * All reap bookkeeping goes through here because [reap] is touched from two places on + * different threads - the trim listener (main) and [fetch]/[warm] (the caller's dispatcher). + * Making every mutation main-thread-only removes the race by construction; @Volatile would not, + * since scheduling is a read-modify-write. Running inline when already on main keeps the + * cancel-then-navigate ordering inside [fetch]'s `Dispatchers.Main` block intact. + */ + private fun onMain(block: () -> Unit) { + if (Looper.myLooper() == Looper.getMainLooper()) block() else main.post(block) + } + + /** + * Free the WebView after a quiet period, like the other four fetchers (issue #182) - this one + * never had it, so a single search pinned a Chromium renderer for the WHOLE session and only + * renderer death or a trim could clear it. Measured on the M5: after a search the shared + * sandboxed renderer sat at 327 MB PSS, in a SEPARATE process, so it never showed up in the + * app's own PSS and every issue-#83 measurement missed it. + * + * [delayMs] differs by caller on purpose. A real fetch uses the siblings' [REAP_IDLE_MS]; a + * speculative [warm] uses the longer [WARM_REAP_IDLE_MS], because the warm exists precisely so + * a later place tap skips the cold start, and reaping it at 120 s would undo it for the ordinary + * search-then-browse-then-tap flow. Longer, but still bounded: the point is that it cannot be + * session-long. + */ + private fun scheduleReap(delayMs: Long) = onMain { + reap?.let(main::removeCallbacks) + val r = Runnable { reap = null; reapNow() } + reap = r + main.postDelayed(r, delayMs) + } + + private fun cancelReap() = onMain { + reap?.let(main::removeCallbacks) + reap = null + } + /** Destroy the WebView immediately. Main thread only (WebView requirement). The next - * [warm]/fetch rebuilds it via `ensureWebView`, exactly as after a renderer death. */ + * [warm]/fetch rebuilds it via `ensureWebView`, exactly as after a renderer death. + * + * Drains [pending] for the same reason [rendererGone] does: the injected scraper dies with the + * view, so nothing will ever complete those deferreds. Without this a reap landing mid-fetch + * leaves the fetch parked in `deferred.await()` for the full [TOTAL_TIMEOUT_MS] while it HOLDS + * [mutex], stalling every queued gallery behind it. An empty result is the documented + * best-effort failure mode; a 40 s hang is not. */ private fun reapNow() { val wv = webView ?: return webView = null warmed = false runCatching { wv.loadUrl("about:blank"); wv.destroy() } + val stranded = pending.keys.toList() + stranded.forEach { id -> pending.remove(id)?.complete("") } + // The renderer is shared and lives in ANOTHER process, so its cost is invisible in this + // app's PSS - without a log there is no way to tell an idle reap from an OS trim killing + // the renderer, which made the first attempt to verify this unfalsifiable. + android.util.Log.i("WebPhotoFetcher", "photo WebView reaped (stranded fetches: ${stranded.size})") } // featureId → its scraped gallery. Re-tapping a place (or bouncing back from directions) then @@ -118,6 +168,10 @@ class WebPhotoFetcher @Inject constructor( } } wv.loadUrl("https://www.google.com/maps?hl=en") + // Bound the speculative warm. Reached only when THIS call created the view (the + // re-check above returns early when a fetch already owns it, and that fetch arms its + // own reap in `finally`), so this never shortens a real fetch's window. + scheduleReap(WARM_REAP_IDLE_MS) } } } @@ -144,6 +198,7 @@ class WebPhotoFetcher @Inject constructor( val cid = cidOf(featureId) ?: return emptyList() synchronized(cache) { cache[featureId] }?.let { return it } // instant on revisit - skip the scrape return mutex.withLock { + cancelReap() // a reap mid-scrape would destroy the view this fetch is about to drive val id = "p" + seq.incrementAndGet() val deferred = CompletableDeferred() pending[id] = deferred @@ -192,6 +247,7 @@ class WebPhotoFetcher @Inject constructor( } finally { pending.remove(id) partials.remove(id) + scheduleReap(REAP_IDLE_MS) // start the quiet period from the END of the scrape } val out = raw?.let { parseLines(it) } ?: emptyList() if (out.isNotEmpty()) synchronized(cache) { cache[featureId] = out } // cache only real results @@ -313,5 +369,11 @@ class WebPhotoFetcher @Inject constructor( // Offscreen viewport so the virtualized category grids render a full batch (not ~1 tile). const val WV_WIDTH = 1200 const val WV_HEIGHT = 3200 + // Destroy the idle WebView after this quiet period, same value as the other four fetchers. + const val REAP_IDLE_MS = 120_000L + // The speculative warm gets a longer leash: it is spent so the first place tap is instant, + // and 120 s would expire during an ordinary browse and waste the warm entirely. Still + // bounded, which is the whole point - before this the warm was held for the session. + const val WARM_REAP_IDLE_MS = 300_000L } } diff --git a/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt b/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt index 90e77ae4..b220ad06 100644 --- a/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt +++ b/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt @@ -54,14 +54,24 @@ class WebPopularTimesFetcher @Inject constructor( @Volatile private var warm: CompletableDeferred? = null private var reap: Runnable? = null + /** Run [block] on the main thread, inline when already there. [reap] is touched from the trim + * listener (main) and from [fetch]/[prewarm] (the caller's dispatcher), and scheduling is a + * read-modify-write that @Volatile would not make safe. */ + private fun onMain(block: () -> Unit) { + if (Looper.myLooper() == Looper.getMainLooper()) block() else main.post(block) + } + /** Free the WebView after a quiet period (issue #182): the warm session pins a full * maps.google.com page for the rest of the session otherwise. The next fetch after a - * reap re-warms (google.com -> maps), a one-off few-second cost after minutes idle. */ - private fun scheduleReap() { + * reap re-warms (google.com -> maps), a one-off few-second cost after minutes idle. + * + * [delayMs] is longer for a speculative [prewarm] than for a real fetch - see + * [WARM_REAP_IDLE_MS]. */ + private fun scheduleReap(delayMs: Long = REAP_IDLE_MS) = onMain { reap?.let(main::removeCallbacks) - val r = Runnable { reapNow() } + val r = Runnable { reap = null; reapNow() } reap = r - main.postDelayed(r, REAP_IDLE_MS) + main.postDelayed(r, delayMs) } /** Destroy the WebView immediately. Must run on the main thread (WebView requirement). */ @@ -80,7 +90,7 @@ class WebPopularTimesFetcher @Inject constructor( } } - private fun cancelReap() { + private fun cancelReap() = onMain { reap?.let(main::removeCallbacks) reap = null } @@ -136,6 +146,12 @@ class WebPopularTimesFetcher @Inject constructor( * idempotent (a warm already in progress is awaited, not restarted). */ suspend fun prewarm() { runCatching { withTimeoutOrNull(MAX_WARM_MS + 2_000L) { ensureWarm() } } + // Bound the speculative warm. Without this a prewarm-only WebView was never reaped at all - + // scheduleReap() is called only from fetch() - so one search pinned the SHARED Chromium + // renderer for the whole session. Device-verified: reaping WebPhotoFetcher's view alone left + // the renderer alive at ~200 MB PSS because this fetcher still held one. A real fetch + // cancels this and re-arms the shorter window in its own finally. + scheduleReap(WARM_REAP_IDLE_MS) } private suspend fun ensureWarm() = withContext(Dispatchers.Main) { @@ -219,6 +235,10 @@ class WebPopularTimesFetcher @Inject constructor( private companion object { const val TOTAL_TIMEOUT_MS = 22_000L const val REAP_IDLE_MS = 120_000L // destroy the idle WebView after this quiet period (issue #182) + // A speculative prewarm gets a longer leash than a real fetch: it is spent so the first + // place tap is fast, and 120 s would expire during an ordinary browse and waste it. Bounded + // is the point - before this a prewarm-only view was never reaped at all. + const val WARM_REAP_IDLE_MS = 300_000L const val SETTLE_MS = 1_200L const val MAX_WARM_MS = 9_000L } diff --git a/docs/FEATURES.md b/docs/FEATURES.md index be37369f..5782ceea 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -284,7 +284,7 @@ Status legend: [x] done · [~] partial / in progress · [ ] planned - [x] CI builds, tests, signs and publishes a normal release `v0.0.` with debug and release APKs; no prerelease channel, tracked by Obtainium and the updater with zero config. - [x] **Opt-in diagnostics/debug export** (Settings → Diagnostics, off by default) - a local-only event log exportable to JSON via the share sheet, never auto-uploaded, wiped when turned off, in-memory only. - [x] **Crash/ANR/jank capture, all local** - an uncaught-exception handler persists stack traces + breadcrumbs to disk for export, ApplicationExitInfo harvests ANR/native/low-memory kills, a debug ANR watchdog and StrictMode flag stalls and main-thread I/O; captured even with diagnostics off, never auto-sent. -- [x] **Memory use cut across the board** (issue #83) - the on-device speech model costs ~267 MB while loaded and used to stay resident for the entire session; it is now dropped after two minutes unused and rebuilt on next use, which reclaims ~101 MB on ANY phone at no visible cost. Every large or native allocation also releases when the system reports memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and gave back nothing when asked. Measured on a 2.9 GB test device: idle memory down 29%, post-trim down 38%, native heap down 57%. Freed memory is also handed back to the operating system rather than kept on the allocator's free lists, which a release alone does not do (a measured ~3 MB more returned per memory warning, at no cost to responsiveness). +- [x] **Memory use cut across the board** (issue #83) - the on-device speech model costs ~267 MB while loaded and used to stay resident for the entire session; it is now dropped after two minutes unused and rebuilt on next use, which reclaims ~101 MB on ANY phone at no visible cost. Every large or native allocation also releases when the system reports memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and gave back nothing when asked. Measured on a 2.9 GB test device: idle memory down 29%, post-trim down 38%, native heap down 57%. Freed memory is also handed back to the operating system rather than kept on the allocator's free lists, which a release alone does not do (a measured ~3 MB more returned per memory warning, at no cost to responsiveness). The two hidden browser views that a search warms up in advance are now released after five idle minutes instead of being held for the whole session: on a 2.9 GB test device that returns about 390 MB to the app and shuts down a separate ~305 MB browser renderer process, with the views rebuilt automatically the next time a place is opened. - [x] **Low-RAM device adaptation** (issue #83) - additionally, the app detects a memory-constrained phone (`isLowRamDevice` or a heap class of 128 MB or less) and adapts: the on-device speech model is not preloaded at startup, the image cache is capped at 16 MB instead of 48 MB, WebView renderers are not warmed speculatively, and the ambient POI fan-out drops from 15 category terms to 8 with a smaller result pool. Independently, every large or native allocation now releases on OS memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and returned nothing when the OS asked. Measured on a 2.9 GB test device: peak memory down 30%, post-trim memory down 38%, native heap down 57%, cold start ~480 ms faster. x86 and x86_64 native libraries are no longer shipped (no target phone can run them), cutting 23 MB of dead weight from every install. - [x] **Trip recording + replay** (Settings → "Save my trips", off by default, separate opt-in) - records each drive's GPS trace to a local file replayable on the map at 3x through the real nav pipeline; saved on arrival, listed with Replay/Share/Delete, Share exporting the raw CSV; replay auto-routes to the destination. - [x] **Simulate driving (demo mode)** (Settings → Navigation, off by default) - Start drives any planned route as a synthetic GPS trace through the live-nav loop so nav runs anywhere for demos and screenshots; End stops it; turn off to navigate for real. From 6e49208d89b8a5cf5190a5abd5636b19ab320975 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Tue, 21 Jul 2026 00:27:49 -0400 Subject: [PATCH 04/21] Stop the ASR reaper freeing the speech model under a running decode The inFlight guard added with the idle reaper did not actually close the window it documents. release() read the counter BEFORE taking loadLock, and ensureRecognizer() handed the native pointer out on a lock-free fast path (`recognizer?.let { if (loadedLang == lang) return it }`) before taking the lock at all, so the two never ordered against each other: reaper release(): inFlight.get() -> 0, guard passes, lock not yet held user listen(): inFlight.incrementAndGet() -> 1 user ensureRecognizer() fast path returns the live pointer user decode starts on it reaper synchronized(loadLock) { recognizer = null; r.release() } user rec.decode(stream) on freed C++ memory That is a SIGSEGV inside libsherpa-onnx-jni, not an exception - the runCatching around the decode cannot catch a native abort, so the whole process dies. An AtomicInteger does not make a check-then-act atomic. The window was not narrow either. Any thread holding loadLock for the ~1 s native load parks release() between its check and its free for that entire time, and warmUp() takes that lock at startup on every non-low-RAM device. Fix: leases, mutated only under loadLock. acquireRecognizer() loads and takes a lease under one lock; release() checks the count and frees under that same lock, so a listen cannot start between the two. The lock-free fast path is gone - an uncontended lock per listen is nothing next to a 15 s utterance. listen() takes the lease and holds it for the whole utterance (recording included) and gives it back in a finally, which keeps it exception-safe against listenInner's many early returns; listenInner now receives the recognizer instead of fetching it. releaseLease() deliberately does NOT take the lock: it runs only once the decode is done with the pointer, so a racing release() can at worst read the pre-decrement value and conservatively decline, and taking the lock there would park the end of every utterance behind an unrelated load. loadLock becomes a ReentrantLock so release() can tryLock instead of blocking. It is called from onTrimMemory on the MAIN thread, and blocking the UI thread for a whole model load to reclaim memory is a bad trade when the idle reaper or the next trim retries anyway. deleteAsrModel passes wait = true, since there the user asked for it and a brief wait is correct. Device-verified on the M5, and the race window is real rather than theoretical: hammering `am send-trim-memory RUNNING_CRITICAL` across startup logged "release skipped, model load in progress" 5 times in one run, i.e. 5 trims landed while the load held the lock - each of which the old code would have used to block the main thread, with the check-then-act live. Across that run and a 6-round force-stop/trim-storm: 0 crashes, 0 SIGSEGV, process alive every round. Release path still works with the model loaded (scudo:secondary 110,564 KB -> 8,031 KB on a severe trim, "recognizer released" logged), the model reloads cleanly afterwards (108,381 KB) and is not quarantined, and the map renders. RUNNING_CRITICAL is the useful level for this test: it is isSevere AND the OS accepts it on a foreground process, so it reaches the load window without needing HOME first. Verified: :app:detekt and :core:detekt 0 smells, :core:test green, audit_deadcode.sh PASS, standardDebug builds, installs and runs, screenshot. --- AGENTS.md | 18 ++ .../main/java/app/vela/ui/map/MapViewModel.kt | 2 +- .../java/app/vela/voice/WhisperRecognizer.kt | 223 ++++++++++++------ 3 files changed, 171 insertions(+), 72 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 61902f49..3d5445f9 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -743,6 +743,24 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): not mean your timer fired. The first attempt to verify this was unfalsifiable for exactly that reason: the renderer vanished at t+60 s and the logs showed `dispatch level=15`/`40`, i.e. a real trim, not the reaper. Log the reap, then assert the log AND assert no severe trim fired. + - **Freeing a native model needs a LEASE, not an atomic counter checked outside the lock.** + `WhisperRecognizer` guards the recognizer with `leases`, mutated ONLY under `loadLock`, and + `release()` checks the count and frees inside that same lock. The first version of this checked + an `AtomicInteger` before taking the lock while `ensureRecognizer()` handed the pointer out on a + lock-free fast path, which is a check-then-act: the reaper reads 0, a mic tap increments and + takes the pointer, the reaper frees it under the running decode. `runCatching` around the decode + CANNOT save you - `OfflineRecognizer.release()` frees C++ memory and the result is a SIGSEGV in + `libsherpa-onnx-jni` that takes the process down. **An atomic counter does not make a + check-then-act atomic.** Take the lease and the pointer under one lock, or do not take either. + - **`release()` must not block the main thread, so it `tryLock`s.** It is called from + `onTrimMemory` on the main thread, and `loadLock` is held across the ~1 s native model load, so + a blocking acquire stalls the UI thread for that whole load just to reclaim memory the idle + reaper would reclaim anyway. Skipping is safe; the next trim or the reaper retries. `Remove + model` passes `wait = true` because there the user asked for it. Device-verified that this + window is REAL and not theoretical: hammering `am send-trim-memory RUNNING_CRITICAL` + across startup logged `release skipped, model load in progress` 5 times in one run. + - `RUNNING_CRITICAL` is the level to use for this - it is `isSevere` AND the OS accepts it on a + FOREGROUND process, so it exercises the load window without needing HOME first. - **A quarantined ASR model makes `warmUp()` a silent no-op.** `WhisperRecognizer.isInstalled()` is `AsrModel.isInstalled() && !asr_model_bad`, and the corrupt-model quarantine (`asr_model_bad` in `vela_settings`) is only ever lifted by the installer's `clearQuarantine()`. Side-loading the diff --git a/app/src/main/java/app/vela/ui/map/MapViewModel.kt b/app/src/main/java/app/vela/ui/map/MapViewModel.kt index 134eabf9..b03ace91 100644 --- a/app/src/main/java/app/vela/ui/map/MapViewModel.kt +++ b/app/src/main/java/app/vela/ui/map/MapViewModel.kt @@ -4369,7 +4369,7 @@ class MapViewModel @Inject constructor( // Free the loaded model BEFORE removing its files. Deleting the directory alone left the // native recognizer resident for the rest of the process (~267 MB measured, issue #83), so // "Remove" reclaimed disk but no memory at all. - whisperRecognizer.release() + whisperRecognizer.release(wait = true) // deliberate user action: worth waiting out a load app.vela.voice.AsrModel.dir(appContext).deleteRecursively() _state.update { it.copy(asrInstalled = false) } } diff --git a/app/src/main/java/app/vela/voice/WhisperRecognizer.kt b/app/src/main/java/app/vela/voice/WhisperRecognizer.kt index 1edf95f1..2e1e82e1 100644 --- a/app/src/main/java/app/vela/voice/WhisperRecognizer.kt +++ b/app/src/main/java/app/vela/voice/WhisperRecognizer.kt @@ -44,15 +44,27 @@ import kotlin.math.sqrt class WhisperRecognizer @Inject constructor( @ApplicationContext private val context: Context, ) { - private val loadLock = Any() + /** Guards [recognizer]/[loadedLang] AND [leases]. A ReentrantLock rather than `synchronized` + * so [release] can `tryLock` instead of blocking - see there. */ + private val loadLock = java.util.concurrent.locks.ReentrantLock() @Volatile private var recognizer: OfflineRecognizer? = null @Volatile private var loadedLang: String? = null - /** Non-zero while a [listen] is inside the native recognizer. [release] refuses to free the - * model while this is set: `OfflineRecognizer.release()` frees C++ memory that an in-flight - * decode is still reading, which is a use-after-free that takes the process down rather than - * throwing. A trim arriving mid-utterance simply keeps the model until the utterance ends. */ - private val inFlight = java.util.concurrent.atomic.AtomicInteger(0) + /** + * Outstanding leases on the loaded recognizer: non-zero while a [listen] holds a pointer to it. + * [release] refuses to free the model while this is set, because `OfflineRecognizer.release()` + * frees C++ memory an in-flight decode is still reading - a use-after-free that takes the + * process down rather than throwing, so `runCatching` around the decode cannot save it. + * + * **Only ever mutated while holding [loadLock]**, in [acquireRecognizer]/[releaseLease]. That is + * the whole point. It used to be incremented in [listen] with no lock while [release] read it + * with no lock, which is a check-then-act with a real window: the reaper could evaluate the + * count as 0, a mic tap could then increment it and take the pointer off `ensureRecognizer`'s + * lock-free fast path, and the reaper would go on to free the model under the running decode. + * The window was not small either - any thread holding [loadLock] for a ~1 s model load parks + * `release()` between its check and the free for that whole time. + */ + private val leases = java.util.concurrent.atomic.AtomicInteger(0) /** Idle-reap timer, same idea as the web fetchers' `REAP_IDLE_MS` (issue #182). One daemon * thread, shared, created lazily so a device that never loads the model never starts it. */ @@ -92,23 +104,67 @@ class WhisperRecognizer @Inject constructor( /** * Free the native recognizer. Safe to call any time: no-op when nothing is loaded, and declines - * while a listen is in flight (see [inFlight]). The next [listen]/[warmUp] rebuilds it. + * while a listen holds a lease (see [leases]). The next [listen]/[warmUp] rebuilds it. + * + * The lease check happens INSIDE [loadLock], together with the free, so a listen cannot start + * between the two. That is what makes this not a use-after-free. + * + * [wait] controls what happens when the lock is already held, which means a ~1 s native model + * load is in progress. The default does NOT block: the trim path runs on the MAIN thread from + * `Application.onTrimMemory`, and stalling the UI thread for a whole model load to reclaim + * memory is a bad trade when the idle reaper or the next trim will retry anyway. `Remove model` + * passes true, because there the user asked for it and a brief wait is correct. */ - fun release() { - if (inFlight.get() > 0) { - Timber.tag(TAG).i("release skipped, listen in flight") + fun release(wait: Boolean = false) { + if (wait) loadLock.lock() else if (!loadLock.tryLock()) { + Timber.tag(TAG).i("release skipped, model load in progress") return } - synchronized(loadLock) { + try { + if (leases.get() > 0) { + Timber.tag(TAG).i("release skipped, listen in flight") + return + } val r = recognizer ?: return recognizer = null loadedLang = null runCatching { r.release() } .onFailure { Timber.tag(TAG).w(it, "recognizer release failed") } Timber.tag(TAG).i("recognizer released") + } finally { + loadLock.unlock() } } + /** + * Load if needed and take a LEASE, both under [loadLock]. Pair with [releaseLease] in a + * `finally`. Returns null when the model is absent or the native load failed, in which case no + * lease is taken. + * + * Callers must not hold the returned pointer past [releaseLease]: the lease is the only thing + * stopping [release] from freeing it. + */ + private fun acquireRecognizer(): OfflineRecognizer? { + loadLock.lock() + try { + val r = ensureRecognizerLocked() ?: return null + leases.incrementAndGet() + return r + } finally { + loadLock.unlock() + } + } + + /** + * Give back a lease taken by [acquireRecognizer]. Deliberately does NOT take [loadLock]: the + * decrement happens only once the decode is finished with the pointer, so the worst a racing + * [release] can do is read the pre-decrement value and conservatively decline. Taking the lock + * here would instead park the end of every utterance behind an unrelated model load. + */ + private fun releaseLease() { + leases.decrementAndGet() + } + private val audioManager by lazy { context.getSystemService(Context.AUDIO_SERVICE) as? AudioManager } @Volatile private var focusRequest: AudioFocusRequest? = null @@ -205,63 +261,78 @@ class WhisperRecognizer @Inject constructor( * null if the model isn't installed or the native load fails - callers then fall back to the * provider intent or hide the mic. */ private fun ensureRecognizer(): OfflineRecognizer? { + loadLock.lock() + try { + return ensureRecognizerLocked() + } finally { + loadLock.unlock() + } + } + + /** + * The body of [ensureRecognizer]. **Caller must hold [loadLock].** + * + * There is deliberately NO lock-free fast path here any more. The old one + * (`recognizer?.let { if (loadedLang == lang) return it }` before the lock) is what let a decode + * obtain the native pointer while [release] was between its lease check and its free. Taking the + * lock on every acquire costs an uncontended lock per listen, which is nothing next to a 15 s + * utterance, and it is what makes the lease in [acquireRecognizer] atomic. + */ + private fun ensureRecognizerLocked(): OfflineRecognizer? { val lang = whisperLang() - recognizer?.let { if (loadedLang == lang) return it } - synchronized(loadLock) { - recognizer?.let { if (loadedLang == lang) return it else runCatching { it.release() } } - recognizer = null - if (!AsrModel.isInstalled(context)) return null + recognizer?.let { if (loadedLang == lang) return it else runCatching { it.release() } } + recognizer = null + if (!AsrModel.isInstalled(context)) return null - // CRASH SENTINEL around the native load. sherpa-onnx parses the .onnx files in C++, and a - // TRUNCATED-but-non-empty model segfaults inside libsherpa-onnx-jni rather than throwing - - // `runCatching` cannot catch a native abort, it takes the whole process down. That is - // reachable in the real world: the installer deletes destDir BEFORE moving staging into - // place, so a copy that stops partway (storage full on a flip phone, process killed) can - // leave a short file that isInstalled()'s present-and-non-empty test happily accepts. - // Because warmUp() runs at STARTUP, the result was an unrecoverable crash loop - the user - // cannot even reach Settings to delete the model, since the app dies before any UI. - // So: mark before the load, clear after. Finding the mark still set means the previous - // attempt never returned - quarantine the model, tell the UI it is not installed, and let - // the app start. Same idiom as the map's two-crash sentinel in VelaMapView. - val prefs = context.getSharedPreferences("vela_settings", Context.MODE_PRIVATE) - if (prefs.getBoolean(KEY_LOAD_INFLIGHT, false)) { - Timber.tag(TAG).e("previous ASR load never returned (native crash) - quarantining the model") - prefs.edit().putBoolean(KEY_LOAD_INFLIGHT, false).putBoolean(KEY_MODEL_BAD, true).apply() - runCatching { AsrModel.dir(context).deleteRecursively() } - return null - } - if (prefs.getBoolean(KEY_MODEL_BAD, false)) return null - prefs.edit().putBoolean(KEY_LOAD_INFLIGHT, true).apply() + // CRASH SENTINEL around the native load. sherpa-onnx parses the .onnx files in C++, and a + // TRUNCATED-but-non-empty model segfaults inside libsherpa-onnx-jni rather than throwing - + // `runCatching` cannot catch a native abort, it takes the whole process down. That is + // reachable in the real world: the installer deletes destDir BEFORE moving staging into + // place, so a copy that stops partway (storage full on a flip phone, process killed) can + // leave a short file that isInstalled()'s present-and-non-empty test happily accepts. + // Because warmUp() runs at STARTUP, the result was an unrecoverable crash loop - the user + // cannot even reach Settings to delete the model, since the app dies before any UI. + // So: mark before the load, clear after. Finding the mark still set means the previous + // attempt never returned - quarantine the model, tell the UI it is not installed, and let + // the app start. Same idiom as the map's two-crash sentinel in VelaMapView. + val prefs = context.getSharedPreferences("vela_settings", Context.MODE_PRIVATE) + if (prefs.getBoolean(KEY_LOAD_INFLIGHT, false)) { + Timber.tag(TAG).e("previous ASR load never returned (native crash) - quarantining the model") + prefs.edit().putBoolean(KEY_LOAD_INFLIGHT, false).putBoolean(KEY_MODEL_BAD, true).apply() + runCatching { AsrModel.dir(context).deleteRecursively() } + return null + } + if (prefs.getBoolean(KEY_MODEL_BAD, false)) return null + prefs.edit().putBoolean(KEY_LOAD_INFLIGHT, true).apply() - val dir = AsrModel.dir(context) - val r = runCatching { - OfflineRecognizer( - config = OfflineRecognizerConfig( - featConfig = FeatureConfig(sampleRate = SAMPLE_RATE, featureDim = 80), - modelConfig = OfflineModelConfig( - whisper = OfflineWhisperModelConfig( - encoder = File(dir, AsrModel.ENCODER).absolutePath, - decoder = File(dir, AsrModel.DECODER).absolutePath, - language = lang, // pinned to the app language ("" = auto) - task = "transcribe", - tailPaddings = -1, - ), - tokens = File(dir, AsrModel.TOKENS).absolutePath, - numThreads = 2, - modelType = "whisper", + val dir = AsrModel.dir(context) + val r = runCatching { + OfflineRecognizer( + config = OfflineRecognizerConfig( + featConfig = FeatureConfig(sampleRate = SAMPLE_RATE, featureDim = 80), + modelConfig = OfflineModelConfig( + whisper = OfflineWhisperModelConfig( + encoder = File(dir, AsrModel.ENCODER).absolutePath, + decoder = File(dir, AsrModel.DECODER).absolutePath, + language = lang, // pinned to the app language ("" = auto) + task = "transcribe", + tailPaddings = -1, ), + tokens = File(dir, AsrModel.TOKENS).absolutePath, + numThreads = 2, + modelType = "whisper", ), - ) - }.getOrNull() - // The load RETURNED (success or a catchable failure), so the process survived it: clear - // the sentinel. Only a native abort leaves it set, which is exactly the case we want the - // next launch to notice. - prefs.edit().putBoolean(KEY_LOAD_INFLIGHT, false).apply() - recognizer = r - loadedLang = lang - if (r != null) armIdleReap() // start the quiet-period countdown from the load - return r - } + ), + ) + }.getOrNull() + // The load RETURNED (success or a catchable failure), so the process survived it: clear + // the sentinel. Only a native abort leaves it set, which is exactly the case we want the + // next launch to notice. + prefs.edit().putBoolean(KEY_LOAD_INFLIGHT, false).apply() + recognizer = r + loadedLang = lang + if (r != null) armIdleReap() // start the quiet-period countdown from the load + return r } /** @@ -273,25 +344,37 @@ class WhisperRecognizer @Inject constructor( /** Listen, transcribe, and say WHY when it does not work - see [VoiceResult]. Every failure exit * logs under `VELAASR` so a tester's logcat names the cause without another round-trip. * - * Thin wrapper over [listenInner] that marks the recognizer busy, so a memory trim arriving - * mid-utterance cannot free the native model out from under the decode (see [inFlight]). The - * inner function keeps its many early returns; this keeps the guard exception-safe. */ + * Thin wrapper over [listenInner] that holds a LEASE on the recognizer for the whole utterance, + * so a memory trim or the idle reaper arriving mid-utterance cannot free the native model out + * from under the decode (see [leases] and [acquireRecognizer]). Taking the lease out here, not + * inside the inner function, is what makes it exception-safe against that function's many + * early returns. */ suspend fun listen( onLevel: (Float) -> Unit, onListening: () -> Unit, cancelled: () -> Boolean, ): VoiceResult { - inFlight.incrementAndGet() + // Acquire (and load) OFF the main thread: this can be a ~1 s native load, and callers reach + // listen() from a UI coroutine. The lease is taken here rather than inside listenInner so + // that the many early returns in there cannot leak it - the finally below always gives it + // back, and it covers recording as well as decoding. reapTask?.cancel(false) // never reap mid-utterance + val rec = withContext(Dispatchers.Default) { acquireRecognizer() } + ?: run { + Timber.tag(TAG).e("listen failed: MODEL (model absent or native load failed)") + armIdleReap() + return VoiceResult.Failed(VoiceResult.Reason.MODEL, "model absent or native load failed") + } try { - return listenInner(onLevel, onListening, cancelled) + return listenInner(rec, onLevel, onListening, cancelled) } finally { - inFlight.decrementAndGet() + releaseLease() armIdleReap() // restart the quiet period from the END of this utterance } } private suspend fun listenInner( + rec: OfflineRecognizer, onLevel: (Float) -> Unit, onListening: () -> Unit, cancelled: () -> Boolean, @@ -300,8 +383,6 @@ class WhisperRecognizer @Inject constructor( Timber.tag(TAG).e("listen failed: $reason${detail?.let { " ($it)" } ?: ""}") return VoiceResult.Failed(reason, detail) } - val rec = ensureRecognizer() - ?: return@withContext fail(VoiceResult.Reason.MODEL, "model absent or native load failed") if (!hasMicPermission()) return@withContext fail(VoiceResult.Reason.PERMISSION) val vad = runCatching { From 7dfc2ab9efdd0febb579ce5ea5b3a10ed0e0cb8b Mon Sep 17 00:00:00 2001 From: alltechdev Date: Tue, 21 Jul 2026 01:33:46 -0400 Subject: [PATCH 05/21] Do not lay out the photo scraper's WebView until it actually scrapes WebPhotoFetcher sized its view inside ensureWebView(), i.e. at construction, and warm() goes through ensureWebView(). So a search built a full 1200x3200 composited surface over maps?hl=en - a page with no scrapeable content on it - and held it for the whole 300 s warm window, on a phone whose screen is 480x640. Sizing moves into sizeForScrape(wv), called in fetch() immediately before loadUrl, so the ?cid= page's FIRST layout is already at scrape geometry. The size is unchanged; only when it is applied changes. Matched A/B, same harness, 3 runs per arm, search then browse with no place opened: eager (as before) GL mtrack 448 / 427 / 441 MB TOTAL PSS 866 / 854 / 852 MB deferred GL mtrack 77 / 71 / 72 MB TOTAL PSS 485 / 460 / 459 MB -365 MB of GPU memory and -390 MB of PSS, no overlap between the arms, and the scrape is unaffected: 28/28/28 photos on the same place in both arms. This fetcher is the only one that lays out during a warm - WebPopularTimes' prewarm, WebDirections and WebStopDepartures never call measure/layout at all, and WebReviews has no warm. That is why every GL number in this app tracks this one view. Both fetchers now log `scraped N photos/reviews for `, which is what makes a viewport change checkable against scrape QUALITY rather than only memory. A change that halves memory and quietly halves the gallery is a regression no memory metric would show. Shrinking the viewport was tried first and is NOT included. On the photo side 720 px held the count (28 -> 28) on the one place it was A/B'd, but the reviews side returns 0 reviews for every place tried, at 1200 AND at 720 - a pre-existing failure, unrelated to width - so that arm's quality metric was pinned at zero and could not fail. An unfalsifiable check is not evidence, and scrape geometry governs how much of a virtualized grid materializes, so both widths stay stock. Deferring the layout wins the same memory back without changing anything the scraper sees, which is the safer bet by construction rather than by sampling. Note GL mtrack at place-open is bimodal (~490 MB laid out and alive, ~71 MB not), so single readings there are worthless; the warm window is the stable thing to measure, and the numbers above are 3 runs per arm. Verified: :app:detekt and :core:detekt 0 smells, :core:test green, audit_deadcode.sh PASS, standardDebug builds, installs and runs, photo gallery renders with the same photos. --- AGENTS.md | 26 +++++++++++ .../main/java/app/vela/web/WebPhotoFetcher.kt | 44 +++++++++++++++++-- .../java/app/vela/web/WebReviewsFetcher.kt | 14 ++++++ docs/FEATURES.md | 2 +- 4 files changed, 81 insertions(+), 5 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 3d5445f9..8e04a497 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -721,6 +721,32 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): the two scraper WebViews exist. That is the offscreen layouts (`WV_WIDTH`x`WV_HEIGHT`, e.g. 1200x3200) allocating graphics buffers charged to OUR process. Chromium logs `tile memory limits exceeded` at that size. Cutting the offscreen viewport is a real lead. + - **Do not LAY OUT a scraper WebView until it is actually scraping.** `WebPhotoFetcher` sized its + view inside `ensureWebView`, i.e. at construction, and `warm()` goes through `ensureWebView` - so + a speculative warm built a full 1200x3200 composited surface over `maps?hl=en`, a page with zero + scrapeable content, and held it for the entire 300 s warm window on a 480x640 phone. Sizing moved + into `sizeForScrape(wv)`, called immediately before `loadUrl` in `fetch()`. Matched A/B, same + harness, 3 runs per arm, search-then-browse with no place opened: + **`GL mtrack` 448/427/441 MB -> 77/71/72 MB (-365 MB) and TOTAL PSS 866/854/852 MB -> + 485/460/459 MB (-390 MB)**, with the scrape byte-for-byte unaffected (28/28/28 photos on the + same place in both arms). + - This fetcher is the ONLY one that lays out during a warm. `WebPopularTimesFetcher.prewarm`, + `WebDirectionsFetcher` and `WebStopDeparturesFetcher` never call measure/layout at all, and + `WebReviewsFetcher` has no warm. That is why every GL number in this codebase tracks THIS view, + and why a reviews-side change looked like it helped when it could not have. + - The size itself is load-bearing at scrape time - the grids virtualize, so at 0x0 a category tab + renders about one tile and the scrape comes back nearly empty. Size BEFORE `loadUrl` so the + page's first layout is already at scrape geometry. Deferring the layout is safe precisely + because it leaves scrape-time geometry identical; SHRINKING it is not the same bet. + - Gate any viewport change on scrape COUNT, not memory. Both fetchers log + `scraped N photos/reviews for ` for exactly this. A change that halves memory and + quietly halves the gallery is a regression no memory metric shows. + - **A quality metric stuck at zero cannot fail, so it proves nothing.** A 720 px width was tried + and reverted: on the photo side it held (28 -> 28) but on the reviews side the scrape returns 0 + for every place tried at 1200 AND at 720 - a pre-existing failure - so that arm was + unfalsifiable. Check the control can produce a non-zero result before trusting an A/B. + - `GL mtrack` at place-open is BIMODAL (~490 MB laid out and alive, ~71 MB not), so single + readings there mean nothing. Measure the warm window, repeat, and report the spread. - **BOTH speculative warms have to be bounded or neither helps.** The renderer is SHARED by every WebView in the process, so one un-reaped view keeps it alive for all of them. `WebPhotoFetcher` had no idle reaper at all, and `WebPopularTimesFetcher.prewarm()` created a view and never diff --git a/app/src/main/java/app/vela/web/WebPhotoFetcher.kt b/app/src/main/java/app/vela/web/WebPhotoFetcher.kt index eaf6a8dd..55327be7 100644 --- a/app/src/main/java/app/vela/web/WebPhotoFetcher.kt +++ b/app/src/main/java/app/vela/web/WebPhotoFetcher.kt @@ -235,6 +235,9 @@ class WebPhotoFetcher @Inject constructor( // DOM that yields an empty result (safe) instead of the previous place's photos // being returned for THIS featureId (cross-place data). wv.evaluateJavascript("try{document.documentElement.innerHTML=''}catch(e){}", null) + // Size BEFORE navigating, so the ?cid= page's first layout is already at + // scrape geometry and the virtualized grids materialize exactly as before. + sizeForScrape(wv) wv.loadUrl("https://www.google.com/maps?cid=$cid&hl=en&gl=us") main.postDelayed({ if (!ready.isCompleted) ready.complete(Unit) }, MAX_LOAD_MS) ready.await() @@ -250,6 +253,9 @@ class WebPhotoFetcher @Inject constructor( scheduleReap(REAP_IDLE_MS) // start the quiet period from the END of the scrape } val out = raw?.let { parseLines(it) } ?: emptyList() + // Result count, so a change to the offscreen viewport (WV_WIDTH/WV_HEIGHT drive how much + // of the virtualized grid renders) can be A/B'd against scrape QUALITY, not just memory. + android.util.Log.i("WebPhotoFetcher", "scraped ${out.size} photos for $featureId") if (out.isNotEmpty()) synchronized(cache) { cache[featureId] = out } // cache only real results out } @@ -282,15 +288,37 @@ class WebPhotoFetcher @Inject constructor( wv.settings.domStorageEnabled = true wv.settings.userAgentString = VelaConfig.USER_AGENT wv.addJavascriptInterface(Bridge(), "VelaBridge") - // Real offscreen viewport - the category grids are VIRTUALIZED (like the reviews list); at 0×0 a - // category tab renders only ~1 tile, so a tall viewport is what makes each category populate fully. + // NOT laid out here on purpose - see [sizeForScrape]. The WebView stays 0x0 until a real + // fetch needs the grids to materialize. + webView = wv + return wv + } + + /** + * Give the WebView its real offscreen viewport, immediately before a scrape navigates. + * + * The size itself is load-bearing: the category grids are VIRTUALIZED (like the reviews list), + * so at 0x0 a category tab renders only about one tile and the scrape comes back nearly empty. + * That is why the viewport exists at all. + * + * But it only has to exist for a SCRAPE. This used to run in `ensureWebView`, i.e. at + * construction, and [warm] goes through `ensureWebView` - so a speculative warm created a full + * WV_WIDTH x WV_HEIGHT composited surface over `maps?hl=en`, a page with zero scrapeable content, + * and held it for the whole WARM_REAP_IDLE_MS window. Measured on the M5: `GL mtrack` is bimodal, + * ~490 MB with this view laid out and alive versus ~71 MB without, on a 480x640 phone screen. + * This fetcher is the only one that lays out during a warm at all (the other four either never + * call layout or have no warm), which is why every GL number tracked THIS view. + * + * Idempotent, and called before `loadUrl` so the `?cid=` page's FIRST layout is already at scrape + * geometry - that ordering is what keeps the scrape identical. + */ + private fun sizeForScrape(wv: WebView) { + if (wv.width == WV_WIDTH && wv.height == WV_HEIGHT) return wv.measure( android.view.View.MeasureSpec.makeMeasureSpec(WV_WIDTH, android.view.View.MeasureSpec.EXACTLY), android.view.View.MeasureSpec.makeMeasureSpec(WV_HEIGHT, android.view.View.MeasureSpec.EXACTLY), ) wv.layout(0, 0, WV_WIDTH, WV_HEIGHT) - webView = wv - return wv } /** Self-polling DOM scraper: open the gallery, then VISIT EACH CATEGORY TAB (Menu / Food & drink / @@ -367,6 +395,14 @@ class WebPhotoFetcher @Inject constructor( const val SETTLE_MS = 1_200L const val MAX_LOAD_MS = 7_000L // Offscreen viewport so the virtualized category grids render a full batch (not ~1 tile). + // + // Applied by [sizeForScrape] immediately before a scrape navigates, NOT at construction. + // + // These are deliberately UNCHANGED from stock. Shrinking the width to 720 was tried and did + // hold the photo count on the one place it was A/B'd (28 -> 28), but scrape geometry governs + // how much of a virtualized grid materializes, and one place is not enough evidence to risk a + // quieter gallery in a locale or layout nobody sampled. Deferring the layout wins the same + // memory back without changing anything the scraper sees, so the size stays stock. const val WV_WIDTH = 1200 const val WV_HEIGHT = 3200 // Destroy the idle WebView after this quiet period, same value as the other four fetchers. diff --git a/app/src/main/java/app/vela/web/WebReviewsFetcher.kt b/app/src/main/java/app/vela/web/WebReviewsFetcher.kt index 051bb490..875bb97a 100644 --- a/app/src/main/java/app/vela/web/WebReviewsFetcher.kt +++ b/app/src/main/java/app/vela/web/WebReviewsFetcher.kt @@ -121,7 +121,11 @@ class WebReviewsFetcher @Inject constructor( return mutex.withLock { cancelReap() try { + // Result count, so a change to the offscreen viewport (WV_WIDTH/WV_HEIGHT drive how + // much of the virtualized list renders) can be A/B'd against scrape QUALITY, not + // just memory. fetchLocked(cid, onProgress, onPartial) + .also { android.util.Log.i("WebReviewsFetcher", "scraped ${it.size} reviews for $featureId") } } finally { scheduleReap() } @@ -417,6 +421,16 @@ class WebReviewsFetcher @Inject constructor( const val MAX_LOAD_MS = 7_000L // Offscreen viewport for the headless WebView - tall so the virtualized review list renders a // healthy batch per scroll position. + // + // Left at STOCK. A 720 px width was tried and reverted: the reviews scrape returns 0 on the + // test device for every place tried, at 1200 AND at 720 - a pre-existing failure, not caused + // by the width - so the quality metric was pinned at zero and the A/B could not fail. An + // unfalsifiable check is not evidence, and scrape geometry governs how much of the + // virtualized list materializes. Do not narrow this until the scrape works again and + // `scraped N reviews` can be compared on a place with hundreds of them. + // + // Unlike WebPhotoFetcher this fetcher lays out only inside a real fetch (it has no warm), so + // it never contributed to the idle/warm-window cost that the deferred layout there fixes. const val WV_WIDTH = 1200 const val WV_HEIGHT = 6000 } diff --git a/docs/FEATURES.md b/docs/FEATURES.md index 5782ceea..92d7f285 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -284,7 +284,7 @@ Status legend: [x] done · [~] partial / in progress · [ ] planned - [x] CI builds, tests, signs and publishes a normal release `v0.0.` with debug and release APKs; no prerelease channel, tracked by Obtainium and the updater with zero config. - [x] **Opt-in diagnostics/debug export** (Settings → Diagnostics, off by default) - a local-only event log exportable to JSON via the share sheet, never auto-uploaded, wiped when turned off, in-memory only. - [x] **Crash/ANR/jank capture, all local** - an uncaught-exception handler persists stack traces + breadcrumbs to disk for export, ApplicationExitInfo harvests ANR/native/low-memory kills, a debug ANR watchdog and StrictMode flag stalls and main-thread I/O; captured even with diagnostics off, never auto-sent. -- [x] **Memory use cut across the board** (issue #83) - the on-device speech model costs ~267 MB while loaded and used to stay resident for the entire session; it is now dropped after two minutes unused and rebuilt on next use, which reclaims ~101 MB on ANY phone at no visible cost. Every large or native allocation also releases when the system reports memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and gave back nothing when asked. Measured on a 2.9 GB test device: idle memory down 29%, post-trim down 38%, native heap down 57%. Freed memory is also handed back to the operating system rather than kept on the allocator's free lists, which a release alone does not do (a measured ~3 MB more returned per memory warning, at no cost to responsiveness). The two hidden browser views that a search warms up in advance are now released after five idle minutes instead of being held for the whole session: on a 2.9 GB test device that returns about 390 MB to the app and shuts down a separate ~305 MB browser renderer process, with the views rebuilt automatically the next time a place is opened. +- [x] **Memory use cut across the board** (issue #83) - the on-device speech model costs ~267 MB while loaded and used to stay resident for the entire session; it is now dropped after two minutes unused and rebuilt on next use, which reclaims ~101 MB on ANY phone at no visible cost. Every large or native allocation also releases when the system reports memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and gave back nothing when asked. Measured on a 2.9 GB test device: idle memory down 29%, post-trim down 38%, native heap down 57%. Freed memory is also handed back to the operating system rather than kept on the allocator's free lists, which a release alone does not do (a measured ~3 MB more returned per memory warning, at no cost to responsiveness). The two hidden browser views that a search warms up in advance are now released after five idle minutes instead of being held for the whole session: on a 2.9 GB test device that returns about 390 MB to the app and shuts down a separate ~305 MB browser renderer process, with the views rebuilt automatically the next time a place is opened. The hidden view a search warms up in advance is also no longer given its full working size until a place is actually opened, since the warm-up page has nothing to read off it: that cuts about 390 MB while you search and browse, and the gallery is built from exactly the same data as before (verified: the same place returned the same 28 photos before and after). - [x] **Low-RAM device adaptation** (issue #83) - additionally, the app detects a memory-constrained phone (`isLowRamDevice` or a heap class of 128 MB or less) and adapts: the on-device speech model is not preloaded at startup, the image cache is capped at 16 MB instead of 48 MB, WebView renderers are not warmed speculatively, and the ambient POI fan-out drops from 15 category terms to 8 with a smaller result pool. Independently, every large or native allocation now releases on OS memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and returned nothing when the OS asked. Measured on a 2.9 GB test device: peak memory down 30%, post-trim memory down 38%, native heap down 57%, cold start ~480 ms faster. x86 and x86_64 native libraries are no longer shipped (no target phone can run them), cutting 23 MB of dead weight from every install. - [x] **Trip recording + replay** (Settings → "Save my trips", off by default, separate opt-in) - records each drive's GPS trace to a local file replayable on the map at 3x through the real nav pipeline; saved on arrival, listed with Replay/Share/Delete, Share exporting the raw CSV; replay auto-routes to the destination. - [x] **Simulate driving (demo mode)** (Settings → Navigation, off by default) - Start drives any planned route as a synthetic GPS trace through the live-nav loop so nav runs anywhere for demos and screenshots; End stops it; turn off to navigate for real. From d4e6dcb3d5e3134f1f6fdba2f7d6d37d2094b6f8 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Tue, 21 Jul 2026 02:19:07 -0400 Subject: [PATCH 06/21] Drop the scraped page when the scrape ends, not two minutes later After a photo scrape the hidden WebView kept a fully rasterized Google Maps document until the 120 s reap - i.e. through the entire time the user is sitting on the place sheet looking at the photos. fetch()'s finally now navigates it to about:blank. Measured at place-open + 75 s, 3 runs: before GL mtrack ~497 MB TOTAL PSS ~950 MB after GL mtrack 64-70 MB TOTAL PSS 410-508 MB Photo counts unchanged. The WebView, renderer, sockets and cookies all stay alive, so the next place is no colder than before; only the document goes. Resizing is NOT the lever, and that was measured rather than assumed. Shrinking the view back to 0x0 after the scrape was tried first and reclaimed nothing: 494/496/497 MB against a 497/498 MB control. Chromium keeps the tiles it has already rasterized for a live document however small the view gets. The 1200x3200 viewport is 3.84 Mpx = 15 MB of pixels against a measured ~490 MB, so that number was never a viewport buffer - it is a whole composited layer tree against a tile budget. Document lifetime is what moves it. The about:blank navigation then broke the NEXT scrape, and this is the part worth reading. Its onPageFinished fires on the webViewClient the next fetch has just installed, which opens that fetch's load gate before the real page has committed, so the scraper injects into an empty document and returns nothing. onPageFinished now ignores about: URLs; the MAX_LOAD_MS fallback still covers a genuinely stuck load. That bug was only visible when opening a SECOND place: the same place scraped 33 photos when it was the first place opened and 0 when it was the second. A one-place test cannot see a WebView REUSE bug at all, and re-tapping the same place is served from the LRU cache without scraping, so it cannot see one either. Verified after the fix by opening two DIFFERENT places in one session, twice over: 28 then 52, and 28 then 33 against a 33 fresh-open baseline for that second place, no crashes, gallery renders. Verified: :app:detekt and :core:detekt 0 smells, :core:test green, audit_deadcode.sh PASS, standardDebug builds, installs and runs, screenshots of both place sheets with their galleries. --- AGENTS.md | 19 ++++++++++++ .../main/java/app/vela/web/WebPhotoFetcher.kt | 31 +++++++++++++++++++ docs/FEATURES.md | 2 +- 3 files changed, 51 insertions(+), 1 deletion(-) diff --git a/AGENTS.md b/AGENTS.md index 8e04a497..6ab600cb 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -721,6 +721,25 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): the two scraper WebViews exist. That is the offscreen layouts (`WV_WIDTH`x`WV_HEIGHT`, e.g. 1200x3200) allocating graphics buffers charged to OUR process. Chromium logs `tile memory limits exceeded` at that size. Cutting the offscreen viewport is a real lead. + - **Throw the scraped DOCUMENT away when the scrape ends; the viewport is not the lever.** After + a scrape the photo fetcher used to hold a fully rasterized Google Maps page until the 120 s reap, + i.e. through the whole time the user reads the place sheet. `blankAfterScrape()` navigates to + `about:blank` in `fetch()`'s `finally`. Measured at place-open + 75 s: **`GL mtrack` ~497 MB -> + 64-70 MB and TOTAL PSS ~950 MB -> 410-508 MB**, photo counts unchanged. + - Resizing does NOT work, and this was measured before believing it: shrinking the view to 0x0 + after a scrape reclaimed nothing at all (494/496/497 MB against a 497/498 MB control). Chromium + keeps the tiles it has rasterized for a live document regardless of view size. **Document + lifetime is the only lever with leverage here** - the 1200x3200 viewport is 3.84 Mpx = 15 MB at + 4 B/px, yet GL mtrack was ~490 MB, so the number is a whole composited layer tree against a + tile budget, not one viewport buffer. Stop spending device time on geometry. + - **`about:blank` opens the next fetch's load gate early unless you guard for it.** Parking the + view at `about:blank` broke the NEXT scrape: its `onPageFinished` fires on the freshly installed + `webViewClient`, completes the `ready` gate before the real page commits, and the scraper injects + into an empty document. **Caught only by opening a SECOND place** - the same place scraped 33 + photos as the first place opened and 0 as the second. `onPageFinished` now ignores `about:` URLs. + - Test the second place, every time. A one-place test cannot see any bug in WebView REUSE, and + re-tapping the SAME place is served from the LRU cache without scraping at all, so it cannot + see one either. Confirm from the log that two DIFFERENT featureIds actually scraped. - **Do not LAY OUT a scraper WebView until it is actually scraping.** `WebPhotoFetcher` sized its view inside `ensureWebView`, i.e. at construction, and `warm()` goes through `ensureWebView` - so a speculative warm built a full 1200x3200 composited surface over `maps?hl=en`, a page with zero diff --git a/app/src/main/java/app/vela/web/WebPhotoFetcher.kt b/app/src/main/java/app/vela/web/WebPhotoFetcher.kt index 55327be7..13e184aa 100644 --- a/app/src/main/java/app/vela/web/WebPhotoFetcher.kt +++ b/app/src/main/java/app/vela/web/WebPhotoFetcher.kt @@ -220,6 +220,13 @@ class WebPhotoFetcher @Inject constructor( return !(host == "google.com" || host.endsWith(".google.com")) } override fun onPageFinished(view: WebView?, url: String?) { + // IGNORE about:blank. [blankAfterScrape] parks the view there after the + // previous scrape, and that navigation can still be settling when this + // client is installed - its onPageFinished then opens the load gate + // early, the scraper injects into an empty document, and the fetch + // returns 0 photos. Device-caught: the same place scraped 33 photos as + // the first place opened and 0 as the second, until this guard. + if (url == null || url.startsWith("about:")) return main.postDelayed({ if (!ready.isCompleted) ready.complete(Unit) }, SETTLE_MS) } override fun onRenderProcessGone(view: WebView?, detail: RenderProcessGoneDetail?): Boolean { @@ -250,6 +257,7 @@ class WebPhotoFetcher @Inject constructor( } finally { pending.remove(id) partials.remove(id) + blankAfterScrape() // the scraped page is dead weight from here until the next scrape scheduleReap(REAP_IDLE_MS) // start the quiet period from the END of the scrape } val out = raw?.let { parseLines(it) } ?: emptyList() @@ -321,6 +329,29 @@ class WebPhotoFetcher @Inject constructor( wv.layout(0, 0, WV_WIDTH, WV_HEIGHT) } + /** + * Throw the scraped page away the moment the scrape returns, rather than carrying it until the + * reap 120 s later. + * + * The scrape is the only thing that needed the page and it is over - the result is already + * parsed out of the bridge payload. What follows is the user reading the place sheet, which is + * minutes of a fully rasterized Google Maps document serving nobody. + * + * It has to be a NAVIGATION, not a resize. Shrinking the view back to 0x0 was tried first and + * measured to reclaim nothing at all (GL mtrack 494/496/497 MB against a 497/498 MB control): + * Chromium keeps the tiles it has already rasterized for a live document regardless of the + * view's size. Discarding the document is what frees them. + * + * Costs nothing functionally: the next `fetch` blanks the DOM and navigates to its own `?cid=` + * page anyway, so this page was never going to be read again. The WebView itself stays alive, so + * the renderer, HTTP/2 sockets, cookies and JS cache that make the next place fast are all kept - + * which is exactly what destroying it early would have thrown away. + */ + private fun blankAfterScrape() = onMain { + val wv = webView ?: return@onMain + runCatching { wv.loadUrl("about:blank") } + } + /** Self-polling DOM scraper: open the gallery, then VISIT EACH CATEGORY TAB (Menu / Food & drink / * Vibe / By owner) in turn - clicking it, scrolling, and tagging the photos it shows with that * category - then sweep the "All" view for the rest (uncategorized). Bridges "category\turl" lines diff --git a/docs/FEATURES.md b/docs/FEATURES.md index 92d7f285..b5c043c0 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -284,7 +284,7 @@ Status legend: [x] done · [~] partial / in progress · [ ] planned - [x] CI builds, tests, signs and publishes a normal release `v0.0.` with debug and release APKs; no prerelease channel, tracked by Obtainium and the updater with zero config. - [x] **Opt-in diagnostics/debug export** (Settings → Diagnostics, off by default) - a local-only event log exportable to JSON via the share sheet, never auto-uploaded, wiped when turned off, in-memory only. - [x] **Crash/ANR/jank capture, all local** - an uncaught-exception handler persists stack traces + breadcrumbs to disk for export, ApplicationExitInfo harvests ANR/native/low-memory kills, a debug ANR watchdog and StrictMode flag stalls and main-thread I/O; captured even with diagnostics off, never auto-sent. -- [x] **Memory use cut across the board** (issue #83) - the on-device speech model costs ~267 MB while loaded and used to stay resident for the entire session; it is now dropped after two minutes unused and rebuilt on next use, which reclaims ~101 MB on ANY phone at no visible cost. Every large or native allocation also releases when the system reports memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and gave back nothing when asked. Measured on a 2.9 GB test device: idle memory down 29%, post-trim down 38%, native heap down 57%. Freed memory is also handed back to the operating system rather than kept on the allocator's free lists, which a release alone does not do (a measured ~3 MB more returned per memory warning, at no cost to responsiveness). The two hidden browser views that a search warms up in advance are now released after five idle minutes instead of being held for the whole session: on a 2.9 GB test device that returns about 390 MB to the app and shuts down a separate ~305 MB browser renderer process, with the views rebuilt automatically the next time a place is opened. The hidden view a search warms up in advance is also no longer given its full working size until a place is actually opened, since the warm-up page has nothing to read off it: that cuts about 390 MB while you search and browse, and the gallery is built from exactly the same data as before (verified: the same place returned the same 28 photos before and after). +- [x] **Memory use cut across the board** (issue #83) - the on-device speech model costs ~267 MB while loaded and used to stay resident for the entire session; it is now dropped after two minutes unused and rebuilt on next use, which reclaims ~101 MB on ANY phone at no visible cost. Every large or native allocation also releases when the system reports memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and gave back nothing when asked. Measured on a 2.9 GB test device: idle memory down 29%, post-trim down 38%, native heap down 57%. Freed memory is also handed back to the operating system rather than kept on the allocator's free lists, which a release alone does not do (a measured ~3 MB more returned per memory warning, at no cost to responsiveness). The two hidden browser views that a search warms up in advance are now released after five idle minutes instead of being held for the whole session: on a 2.9 GB test device that returns about 390 MB to the app and shuts down a separate ~305 MB browser renderer process, with the views rebuilt automatically the next time a place is opened. The hidden view a search warms up in advance is also no longer given its full working size until a place is actually opened, since the warm-up page has nothing to read off it: that cuts about 390 MB while you search and browse, and the gallery is built from exactly the same data as before (verified: the same place returned the same 28 photos before and after). It also drops the scraped page as soon as the photos have been read instead of holding it for two minutes, which cuts a further ~440 MB while you are reading a place, again with the same photos. - [x] **Low-RAM device adaptation** (issue #83) - additionally, the app detects a memory-constrained phone (`isLowRamDevice` or a heap class of 128 MB or less) and adapts: the on-device speech model is not preloaded at startup, the image cache is capped at 16 MB instead of 48 MB, WebView renderers are not warmed speculatively, and the ambient POI fan-out drops from 15 category terms to 8 with a smaller result pool. Independently, every large or native allocation now releases on OS memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and returned nothing when the OS asked. Measured on a 2.9 GB test device: peak memory down 30%, post-trim memory down 38%, native heap down 57%, cold start ~480 ms faster. x86 and x86_64 native libraries are no longer shipped (no target phone can run them), cutting 23 MB of dead weight from every install. - [x] **Trip recording + replay** (Settings → "Save my trips", off by default, separate opt-in) - records each drive's GPS trace to a local file replayable on the map at 3x through the real nav pipeline; saved on arrival, listed with Replay/Share/Delete, Share exporting the raw CSV; replay auto-routes to the destination. - [x] **Simulate driving (demo mode)** (Settings → Navigation, off by default) - Start drives any planned route as a synthetic GPS trace through the live-nav loop so nav runs anywhere for demos and screenshots; End stops it; turn off to navigate for real. From e5c8200b89c308fb7b7238483c55a3a367fc25c9 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Tue, 21 Jul 2026 02:46:38 -0400 Subject: [PATCH 07/21] Measure the production variant, and correct two claims about the native purge Every memory number in issue #83 and in this branch so far was standardDebug. The `staging` variant exists precisely to avoid that (initWith(release), R8, resources shrunk, non-debuggable, installs side by side as app.vela.staging) and nobody had used it. Production is substantially leaner than debug: state standardDebug standardStaging after a search ~460 MB ~279 MB place open 410-508 MB 335-392 MB after a severe trim - 140-146 MB Code bucket 101 MB 30-45 MB The Code gap is extracted dex and JIT profiles that do not exist in a release build, so roughly 55-70 MB of any debug reading is an artifact. Scrape verified unaffected by R8: 28 photos on staging, same as debug, and the @JavascriptInterface bridge survives minification (onResult is present in the minified dex). Two corrections, both to claims made earlier on this branch. FIRST: "a 442 MB arena holding just 46 MB live" was presented as though the gap were reclaimable. It is not. mallinfo's free figure is address space that scudo has already madvised away; scudo:primary PSS at that same moment was 67 MB. Measured directly by purging during active use with NO listener release - a RUNNING_MODERATE trim, which isSevere excludes - scudo:primary moved 67.1 MB to 64.5 MB. That is 2.6 MB. A periodic idle purge would be worthless and is not worth building; this note exists so nobody builds it. SECOND: the purge was reported as worth ~3 MB from a debug A/B. On production it is ~10 MB. A/B on staging, 3 runs per arm, comparing where scudo:primary SETTLES after a severe trim (the pre-trim value swings 147-382 MB run to run and is useless as a baseline): purge on 52.4/53.1/54.1 MB, purge off 54.4/58.6/76.7 MB. A single trim on production reclaims 133-328 MB, and it would have been easy to credit that to the purge - the first reading looked like 122 MB. Nearly all of it is the registered listeners releasing plus what the platform already does on trim. The control is what separated them. Docs only; no behaviour change. --- AGENTS.md | 28 ++++++++++++++++++++++++++++ 1 file changed, 28 insertions(+) diff --git a/AGENTS.md b/AGENTS.md index 6ab600cb..a0caa45d 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -721,6 +721,34 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): the two scraper WebViews exist. That is the offscreen layouts (`WV_WIDTH`x`WV_HEIGHT`, e.g. 1200x3200) allocating graphics buffers charged to OUR process. Chromium logs `tile memory limits exceeded` at that size. Cutting the offscreen viewport is a real lead. + - **MEASURE THE `staging` VARIANT, NOT `debug`.** `staging` is `initWith(release)` - R8-minified, + resources shrunk, non-debuggable, installs side by side as `app.vela.staging` - so it is the + production memory profile without touching a real release install. Every number in issue #83 was + `standardDebug` and overstates the app substantially: + + | state | standardDebug | standardStaging (production) | + |---|---|---| + | after a search (warm) | ~460 MB | **~279 MB** | + | place open | 410-508 MB | **335-392 MB** | + | after a severe trim | - | **140-146 MB** | + | `Code` bucket | 101 MB | **30-45 MB** | + + The `Code` gap is extracted dex and JIT profiles that simply do not exist in a release build, so + roughly 55-70 MB of any debug reading is an artifact. Check a conclusion against `staging` before + spending effort on it. + - **`mallinfo` "free" is address space, NOT reclaimable resident memory. Do not chase it.** The + arena routinely reports something like 440 MB total against 41 MB live, which looks like ~400 MB + waiting to be reclaimed. It is not: scudo has already madvised those pages away, and + `scudo:primary` PSS at that same moment was only 67 MB. Measured directly by purging during + active use with no listener release (a `RUNNING_MODERATE` trim, which `isSevere` excludes): + `scudo:primary` moved 67.1 -> 64.5 MB, i.e. **2.6 MB**. A periodic idle purge is therefore not + worth building; the earlier framing of that gap as reclaimable was wrong. + - **What the `mallopt` purge is actually worth: ~10 MB, on production.** A/B on `staging`, 3 runs + per arm, comparing where `scudo:primary` SETTLES after a severe trim (the pre-trim value swings + 147-382 MB run to run and is useless): purge on 52.4/53.1/54.1 MB, purge off 54.4/58.6/76.7 MB. + A single trim reclaims 133-328 MB on production, but nearly all of that is the registered + listeners releasing plus what the platform already does on trim - only ~10 MB is the purge. Do + not credit the purge with the whole trim delta; run the control. - **Throw the scraped DOCUMENT away when the scrape ends; the viewport is not the lever.** After a scrape the photo fetcher used to hold a fully rasterized Google Maps page until the 120 s reap, i.e. through the whole time the user reads the place sheet. `blankAfterScrape()` navigates to From bd08647f10a38c391cf7019668c085e9e77907d9 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Tue, 21 Jul 2026 03:04:51 -0400 Subject: [PATCH 08/21] Stop low-RAM phones losing whole POI categories for no memory saving The low-RAM ambient path fetched 8 of the 15 category terms. Both halves of that were wrong. It saved no peak memory. Peak is set by ambientFanout, a Semaphore(4), and every buffer - the response String, the stripped copy, the JsonElement DOM - is allocated INSIDE withPermit. At most 4 of those exist at once however many terms are queued behind them, so 15 to 8 changes how many WAVES the fan-out takes, not what is resident at the peak. The semaphore's own KDoc already said so: "Bounding to 4 caps the peak transient heap with the same final pool." The levers that do move the peak are the permit count and the response size, and the !7i pool halving is the one in use. It is untouched here. And its stated justification was false. It kept school and park on the grounds that only those two lack a second source while the ambient layer is up. Nothing has one then: VelaMapView sets poi_r1/poi_r7/poi_r20 to NONE wholesale on `if (navMode || ambientPois.isNotEmpty())`, not per category. So shopping, services, beauty salon, fast food, gym, bar and pharmacy lost their basemap fallback exactly as school and park would have. Parks at least keep a landuse polygon, so the green area survives without the pin; a gym, a bar or a pharmacy exists ONLY as a POI pin, which makes those the worse things to drop, not the safer ones. A constrained phone was quietly showing a different, poorer map. The observation behind the subset was real - a first 6-term attempt did lose every park and school pin, caught by an A/B screenshot. The generalisation drawn from it was not: that screenshot was evidence about the fan-out, not about school and park being special. Low-RAM devices now fetch the same terms as everyone else, so the category set is identical by construction rather than by sampling. The only remaining low-RAM difference in this path is the smaller !7i result pool, which trims deep-rank results per term without removing any category. Verified: :app:detekt and :core:detekt 0 smells, :core:test green, audit_deadcode.sh PASS, standardDebug builds and installs, and the app launches and renders ambient POIs with debug.vela.lowram forced true - park, shopping and restaurant pins present, no crash. --- AGENTS.md | 21 ++++++++-- .../core/data/google/GoogleMapsDataSource.kt | 39 +++++++++---------- docs/FEATURES.md | 2 +- 3 files changed, 37 insertions(+), 25 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index a0caa45d..c11767a2 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -675,10 +675,23 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): Ask "can this be released when unused?" before "which devices should get less?". - **`MapView.onLowMemory()` must be called.** MapLibre's tile/glyph/sprite caches are native and that is the only way to shrink them; nothing called it before. - - **When you gate a category fan-out down, check what has no SECOND source.** The low-RAM ambient - subset keeps `school` and `park` on purpose: the ambient layer filter-hides the basemap OSM poi - layers at z14+, so those two would vanish entirely. A first 6-term subset did exactly that and - was caught by an A/B screenshot, not by any test. + - **Cutting the ambient TERM COUNT does not cut peak memory, and it silently deletes POIs.** The + low-RAM path briefly fetched 8 of the 15 category terms. Both halves of that were wrong. + - Peak is set by `ambientFanout`, a `Semaphore(4)`, and every buffer (response String, stripped + copy, `JsonElement` DOM) is allocated INSIDE `withPermit`. At most 4 exist at once however many + terms are queued behind them, so 15 -> 8 changes the number of WAVES, not the peak. The + semaphore's own KDoc already said this: "Bounding to 4 caps the peak transient heap with the + same final pool." The levers that DO move the peak are the permit count and the response size + (`!7i`), and only the latter is used. + - Its justification was false. It kept `school` and `park` on the grounds that only they lack a + second source while the ambient layer is up. NOTHING has one then: `VelaMapView` sets + `poi_r1/poi_r7/poi_r20` to `NONE` **wholesale** on `if (navMode || ambientPois.isNotEmpty())`, + not per category. The dropped terms lost their fallback identically. Parks at least keep a + landuse polygon so the green area survives without the pin; a gym, bar or pharmacy exists ONLY + as a pin, making those the worse things to drop, not the safer ones. + - The observation behind it was real (a 6-term subset did lose every park and school pin, caught + by an A/B screenshot). The GENERALISATION drawn from one observation was not. When a screenshot + shows category X vanishing, that is evidence about the fan-out, not about X being special. - **A Kotlin `release()` does NOT give memory back to the KERNEL, only to scudo.** Freeing a model or a WebView returns its pages to the allocator's free lists, where RSS/PSS still count them and lmkd still sees a fat process. `mallopt()` is the only way to hand them on and it is reachable diff --git a/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt b/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt index 6f121bb9..430d8118 100644 --- a/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt +++ b/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt @@ -143,7 +143,7 @@ class GoogleMapsDataSource @Inject constructor( // FAN OUT across category terms + merge: one "places" query is biased to prominent food/ // shops, so it misses whole tiers (a strip mall's plumber, nail salon, IT shop). A handful // of category queries roughly DOUBLES local coverage (live: 22→52 unique within 600 m). - val allTerms = listOf( + val terms = listOf( "places", "restaurants", "coffee", "stores", "shopping", "services", "beauty salon", "fast food", // High-traffic everyday categories the food/shop-biased set above under-returns, so the map // shows a Google-like MIX (a gas station, a gym, a grocer) rather than mostly restaurants. @@ -155,27 +155,26 @@ class GoogleMapsDataSource @Inject constructor( // their prominence low, so they surface in quiet/residential views without crowding businesses. "school", "park", ) - // LOW-RAM: the fan-out is the app's single largest allocation burst. Each term buffers a - // full response String, a stripped copy, and a JsonElement DOM (GoogleResponse.parse), and - // 15 of those run 4-at-a-time per pan - the ~180 MB/12 s churn in AGENTS.md's memory rule. - // Constrained devices fetch an 8-term subset instead of 15. + // LOW-RAM devices fetch the SAME terms as everyone else. An earlier attempt cut this to an + // 8-term subset, and that was wrong twice over (issue #83 follow-up). // - // The subset is NOT just the first N. "school" and "park" are retained DELIBERATELY: while - // the ambient layer is active at z14+ the basemap's OSM poi layers are filter-hidden, so - // those categories have NO second source - dropping them makes parks and schools vanish - // from the map entirely, which is the exact bug the civic/green terms were added to fix - // (see the comment on allTerms above). Device-verified by A/B screenshot: a first attempt - // at a 6-term subset lost every park and school pin on the low-RAM frame. + // It did not save peak memory. Peak is set by [ambientFanout], a Semaphore(4), and every + // buffer - the response String, the stripped copy, the JsonElement DOM - is allocated INSIDE + // `withPermit`. At most 4 of those exist at once no matter how many terms are queued behind + // them, so going 15 -> 8 changes how many WAVES the fan-out takes, not how much is resident + // at the peak. The levers that do move the peak are the permit count and the response size, + // and the `!7i` pool halving below is the one being used. // - // What goes instead are the terms whose places still surface via "places"/"stores" or whose - // absence degrades gracefully: shopping, services, beauty salon, fast food, gym, bar, - // pharmacy. Fewer ambient POIs is a visible trade, and the right one on a phone that - // otherwise OOMs (issue #83). Roomier devices are unaffected. - val terms = if (LowRamMode.enabled) { - listOf("places", "restaurants", "coffee", "stores", "grocery store", "gas station", "school", "park") - } else { - allTerms - } + // And its stated justification was false. It kept "school" and "park" on the grounds that + // only they lack a second source once the ambient layer is active. In fact NOTHING has a + // second source then: VelaMapView sets poi_r1/poi_r7/poi_r20 to NONE wholesale whenever any + // ambient POI exists (`if (navMode || ambientPois.isNotEmpty())`), not per category. So the + // dropped terms - shopping, services, beauty salon, fast food, gym, bar, pharmacy - lost + // their basemap fallback exactly as school and park would have. Parks at least keep their + // landuse polygon, so the green area survives without the pin; a gym, a bar or a pharmacy + // exists ONLY as a POI pin, which makes them the worse thing to drop, not the safer one. + // The A/B screenshot that caught vanishing parks was real; the explanation drawn from it did + // not generalise, and the subset it produced was built on that explanation. suspend fun fetchTerm(term: String): List = ambientFanout.withPermit { runCatching { val pb = SearchPb.build(term, center, cal.searchPb) diff --git a/docs/FEATURES.md b/docs/FEATURES.md index b5c043c0..8814d10c 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -285,7 +285,7 @@ Status legend: [x] done · [~] partial / in progress · [ ] planned - [x] **Opt-in diagnostics/debug export** (Settings → Diagnostics, off by default) - a local-only event log exportable to JSON via the share sheet, never auto-uploaded, wiped when turned off, in-memory only. - [x] **Crash/ANR/jank capture, all local** - an uncaught-exception handler persists stack traces + breadcrumbs to disk for export, ApplicationExitInfo harvests ANR/native/low-memory kills, a debug ANR watchdog and StrictMode flag stalls and main-thread I/O; captured even with diagnostics off, never auto-sent. - [x] **Memory use cut across the board** (issue #83) - the on-device speech model costs ~267 MB while loaded and used to stay resident for the entire session; it is now dropped after two minutes unused and rebuilt on next use, which reclaims ~101 MB on ANY phone at no visible cost. Every large or native allocation also releases when the system reports memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and gave back nothing when asked. Measured on a 2.9 GB test device: idle memory down 29%, post-trim down 38%, native heap down 57%. Freed memory is also handed back to the operating system rather than kept on the allocator's free lists, which a release alone does not do (a measured ~3 MB more returned per memory warning, at no cost to responsiveness). The two hidden browser views that a search warms up in advance are now released after five idle minutes instead of being held for the whole session: on a 2.9 GB test device that returns about 390 MB to the app and shuts down a separate ~305 MB browser renderer process, with the views rebuilt automatically the next time a place is opened. The hidden view a search warms up in advance is also no longer given its full working size until a place is actually opened, since the warm-up page has nothing to read off it: that cuts about 390 MB while you search and browse, and the gallery is built from exactly the same data as before (verified: the same place returned the same 28 photos before and after). It also drops the scraped page as soon as the photos have been read instead of holding it for two minutes, which cuts a further ~440 MB while you are reading a place, again with the same photos. -- [x] **Low-RAM device adaptation** (issue #83) - additionally, the app detects a memory-constrained phone (`isLowRamDevice` or a heap class of 128 MB or less) and adapts: the on-device speech model is not preloaded at startup, the image cache is capped at 16 MB instead of 48 MB, WebView renderers are not warmed speculatively, and the ambient POI fan-out drops from 15 category terms to 8 with a smaller result pool. Independently, every large or native allocation now releases on OS memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and returned nothing when the OS asked. Measured on a 2.9 GB test device: peak memory down 30%, post-trim memory down 38%, native heap down 57%, cold start ~480 ms faster. x86 and x86_64 native libraries are no longer shipped (no target phone can run them), cutting 23 MB of dead weight from every install. +- [x] **Low-RAM device adaptation** (issue #83) - additionally, the app detects a memory-constrained phone (`isLowRamDevice` or a heap class of 128 MB or less) and adapts: the on-device speech model is not preloaded at startup, the image cache is capped at 16 MB instead of 48 MB, WebView renderers are not warmed speculatively, and the ambient POI fan-out asks for a smaller result pool per category. Constrained phones still search every category the roomier ones do, so no kind of place (a gym, a bar, a pharmacy, a school, a park) is ever missing from the map because of the device. Independently, every large or native allocation now releases on OS memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and returned nothing when the OS asked. Measured on a 2.9 GB test device: peak memory down 30%, post-trim memory down 38%, native heap down 57%, cold start ~480 ms faster. x86 and x86_64 native libraries are no longer shipped (no target phone can run them), cutting 23 MB of dead weight from every install. - [x] **Trip recording + replay** (Settings → "Save my trips", off by default, separate opt-in) - records each drive's GPS trace to a local file replayable on the map at 3x through the real nav pipeline; saved on arrival, listed with Replay/Share/Delete, Share exporting the raw CSV; replay auto-routes to the destination. - [x] **Simulate driving (demo mode)** (Settings → Navigation, off by default) - Start drives any planned route as a synthetic GPS trace through the live-nav loop so nav runs anywhere for demos and screenshots; End stops it; turn off to navigate for real. - [x] **Simulate my location (demo mode)** (off by default) - Vela pretends you're at the map centre so the dot, directions origin and recenter read from there; turn off for real GPS. From da37ca7f5dc8b9db97301b346656d028986d01b7 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Tue, 21 Jul 2026 03:27:42 -0400 Subject: [PATCH 09/21] Stop the low-RAM check missing the phones it was written for `heapClassMb in 1..127` had two defects, and both were invisible on any device the dev side owns. It excluded 128, which is the heap class OEMs hand out across 1 GB phones and the low end of 2 GB ones - exactly the class of device issue #83 was filed from. The phone this work was written for could plausibly have matched none of the predicate and received none of the work. And a probe that could not be read returns 0, which also falls outside 1..127, so an unknown device was silently routed down the memory-HUNGRY path. That is the wrong failure direction: landing on the low-RAM path costs a roomy phone about a second on its first mic tap and its first place open, while landing on the normal path can OOM a phone that had no headroom to begin with. The predicate now takes three signals, any one of which is enough: isLowRamDevice (canonical but only Go builds set it), total RAM at or under 2048 MB, and heap class at or under 128 MB. Total RAM is new here and is the signal that actually describes the device - heap class is a Dalvik knob an OEM can set to anything. An unreadable probe never counts as evidence of a roomy device, and two unreadable probes classify as constrained. MemoryInfo.totalMem reports what the OS can hand out rather than the marketing figure, so a nominal 2 GB phone reads about 1900 MB and lands inside the ceiling while a 3 GB phone reads about 2800 MB and does not. Device-confirmed: the M5 reads heapClassMb=256 totalRamMb=2878 and stays lowRam=false, so every measurement recorded in AGENTS.md was taken on the path that still ships to a roomy phone. The decision moved to LowRamMode.classify in :core for one reason: so it can be tested. A predicate that gates every memory adaptation in the app had no test, and both of its bugs were the kind only a device nobody owns would expose. LowRamModeTest pins the boundaries, including 128 itself, the inclusive total-RAM edge, unreadable probes, and the M5's own values. Negative control run, as AGENTS.md requires: restoring the original `isLowRamDevice || heapClassMb in 1..127` makes 4 of the 9 tests fail, including "heap class 128 is low-RAM, the boundary the first version excluded" and "both probes unreadable is treated as constrained, not roomy". The tests fail on the bug they were written for. Verified: :app:detekt and :core:detekt 0 smells, :core:test green (9/9 in the new file), audit_deadcode.sh PASS, standardDebug builds, installs and launches, and the M5 logs lowRam=false unchanged. --- AGENTS.md | 15 ++++ .../main/java/app/vela/ui/MemoryPressure.kt | 47 ++++++++-- .../java/app/vela/core/data/LowRamMode.kt | 46 ++++++++++ .../java/app/vela/core/data/LowRamModeTest.kt | 85 +++++++++++++++++++ docs/FEATURES.md | 2 +- 5 files changed, 186 insertions(+), 9 deletions(-) create mode 100644 core/src/test/java/app/vela/core/data/LowRamModeTest.kt diff --git a/AGENTS.md b/AGENTS.md index c11767a2..bf5088e2 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -647,6 +647,21 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): from a trim would CONSTRUCT it, so the trim would allocate the very thing it is freeing. - `LowRamMode.enabled` (`:core`) is the `:core`-visible mirror, pushed in by `VelaApp` - same seam as `CategoryFilter.enabled`, because `:core` must never read an `:app` holder. + - **The low-RAM predicate lives in `:core` as `LowRamMode.classify`, and it has TESTS.** It has + shipped wrong twice, both times in ways no dev device could reveal: `heapClassMb in 1..127` + excluded 128, the one heap class 1 GB phones and low-end 2 GB phones actually use - so the + device issue #83 was filed from could plausibly have received none of the work - and a 0 from a + failed probe fell out of that range and selected the memory-HUNGRY path for a device we knew + nothing about. It now takes three signals (`isLowRamDevice`, total RAM <= 2048 MB, heap class + <= 128 MB), **treats an unreadable probe as constrained**, and lives in `:core` purely so + `core/src/test/.../LowRamModeTest.kt` can pin the boundaries. Failing toward low-RAM costs a + roomy phone about a second on its first mic tap; failing the other way can OOM a phone with no + headroom. + - Total RAM is the signal that actually describes the device; heap class is a Dalvik knob an OEM + can set to anything. `MemoryInfo.totalMem` reports what the OS can hand out, so a nominal 2 GB + phone reads ~1900 MB and a 3 GB phone ~2800 MB. The M5 reads 2878 MB and stays on the normal + path, which is what keeps every measurement in this file comparable - there is a test asserting + exactly that, so if it ever flips you will be told. - **Verify the low-RAM path or it ships unverified.** Every dev phone we own reports `lowRam=false heapClassMb=256`, so those branches are dead code locally. Debug builds honour `adb shell setprop debug.vela.lowram true` (then relaunch). NB `false` FORCES the normal path, diff --git a/app/src/main/java/app/vela/ui/MemoryPressure.kt b/app/src/main/java/app/vela/ui/MemoryPressure.kt index 98ffd883..88f41751 100644 --- a/app/src/main/java/app/vela/ui/MemoryPressure.kt +++ b/app/src/main/java/app/vela/ui/MemoryPressure.kt @@ -29,26 +29,56 @@ object MemoryPressure { private val listeners = CopyOnWriteArrayList() /** - * True when the OS classes this device as low-RAM (`ActivityManager.isLowRamDevice`) OR its - * heap class is small enough that our normal budgets do not fit. The heap-class arm matters: - * plenty of cheap keypad phones do NOT set the low-RAM system property yet still hand out a - * 96 MB heap class, and those are exactly the phones this work is for. + * True when this device cannot comfortably carry our normal budgets. Three independent signals, + * any of which is enough, because no single one catches the phones this work is for: + * + * - `ActivityManager.isLowRamDevice`, the canonical flag, but only Go-configured builds set it. + * - Total system RAM ([totalRamMb]) at or under `LowRamMode.LOW_TOTAL_RAM_MB`. This is the + * signal that actually describes the device; heap class is a Dalvik tuning knob an OEM can set + * to anything. + * - Heap class ([heapClassMb]) at or under `LowRamMode.LOW_HEAP_CLASS_MB`, which catches an OEM + * that ships plenty of RAM but hands apps a small heap. + * + * The predicate itself is `LowRamMode.classify`, in `:core` so it can be unit-tested. + * + * **When we cannot tell, we assume constrained.** Failing to the low-RAM path costs a roomy + * phone about a second on its first mic tap and its first place open; failing the other way can + * OOM a phone that had no headroom. The old predicate did the opposite: it read + * `heapClassMb in 1..127`, so a failed `ActivityManager` lookup produced 0, fell out of the + * range, and silently selected the memory-hungry path on a device we knew nothing about. */ @Volatile var lowRam: Boolean = false private set - /** The device's normal (non-large) heap class in MB. 0 until [init]. */ + /** The device's normal (non-large) heap class in MB. 0 when it could not be read. */ @Volatile var heapClassMb: Int = 0 private set + /** Total system RAM in MB, as the OS reports it. 0 when it could not be read. */ + @Volatile var totalRamMb: Int = 0 + private set + fun init(context: Context) { val am = context.getSystemService(Context.ACTIVITY_SERVICE) as? ActivityManager heapClassMb = am?.memoryClass ?: 0 + totalRamMb = am?.let { m -> + runCatching { + val mi = ActivityManager.MemoryInfo() + m.getMemoryInfo(mi) + (mi.totalMem / (1024L * 1024L)).toInt() + }.getOrDefault(0) + } ?: 0 val forced = forcedLowRam() - lowRam = forced ?: ((am?.isLowRamDevice == true) || (heapClassMb in 1..127)) + // The decision itself lives in :core so it can be unit-tested; this side only probes. + // am == null means we could not ask at all, which the 0/0 probes already classify as + // constrained, but say it explicitly rather than leaning on that coincidence. + lowRam = forced ?: ( + am == null || + app.vela.core.data.LowRamMode.classify(am.isLowRamDevice, heapClassMb, totalRamMb) + ) Timber.i( - "MemoryPressure init lowRam=%b heapClassMb=%d forced=%s", - lowRam, heapClassMb, forced?.toString() ?: "no", + "MemoryPressure init lowRam=%b heapClassMb=%d totalRamMb=%d forced=%s", + lowRam, heapClassMb, totalRamMb, forced?.toString() ?: "no", ) } @@ -186,6 +216,7 @@ object MemoryPressure { /** Long enough for the main-looper-posted WebView destroys and the Coil clear to have landed. */ private const val PURGE_DELAY_MS = 750L + /** * The app is backgrounded or the OS is genuinely short of memory, so caches that only speed * things up should go. Everything at or above this level is a "drop it" signal. diff --git a/core/src/main/java/app/vela/core/data/LowRamMode.kt b/core/src/main/java/app/vela/core/data/LowRamMode.kt index 56d7272d..5cab1a9c 100644 --- a/core/src/main/java/app/vela/core/data/LowRamMode.kt +++ b/core/src/main/java/app/vela/core/data/LowRamMode.kt @@ -14,4 +14,50 @@ object LowRamMode { /** Set once from `VelaApp.onCreate` after `MemoryPressure.init`. */ @Volatile var enabled: Boolean = false + + /** + * Heap-class ceiling for the low-RAM path, INCLUSIVE. + * + * 128 is the value that matters, and the one the first version of this got wrong by writing + * `in 1..127`: it is the heap class OEMs hand out across 1 GB phones and the low end of 2 GB + * ones, exactly the class of device issue #83 was filed from. Excluding it meant the phone the + * work was written for could plausibly have received none of it. 192 and up stays normal. + */ + const val LOW_HEAP_CLASS_MB = 128 + + /** + * Total-RAM ceiling for the low-RAM path, INCLUSIVE. + * + * `ActivityManager.MemoryInfo.totalMem` reports what the OS can hand out, meaningfully less than + * the marketing figure once the kernel has taken its share: a nominal 2 GB phone reports roughly + * 1900 MB and lands inside this, a 3 GB phone roughly 2800 MB and does not. The M5 dev phone + * reports 2878 MB, so it stays on the normal path and its recorded measurements stay comparable. + */ + const val LOW_TOTAL_RAM_MB = 2048 + + /** + * Decide whether a device is memory-constrained, from probes the `:app` side gathers. + * + * Pure and testable ON PURPOSE. This predicate has already been wrong twice - an exclusive + * `1..127` that skipped the single most important heap class, and a zero-means-roomy fallthrough + * that sent an unknown device down the memory-hungry path - and neither was catchable without a + * device that reproduced it. It is `:core` so it can have unit tests. + * + * Pass 0 for a probe that could not be read; 0 never counts as evidence of a roomy device. + * + * @param isLowRamDevice `ActivityManager.isLowRamDevice`, the canonical flag, which only + * Go-configured builds set. + * @param heapClassMb `ActivityManager.memoryClass`, a Dalvik knob an OEM can set to anything. + * @param totalRamMb total system RAM, the signal that actually describes the device. + */ + fun classify(isLowRamDevice: Boolean, heapClassMb: Int, totalRamMb: Int): Boolean = when { + isLowRamDevice -> true + heapClassMb in 1..LOW_HEAP_CLASS_MB -> true + totalRamMb in 1..LOW_TOTAL_RAM_MB -> true + // Neither probe told us anything. Assume constrained: failing this way costs a roomy phone + // about a second on its first mic tap and first place open, while failing the other way can + // OOM a phone that had no headroom to begin with. + heapClassMb == 0 && totalRamMb == 0 -> true + else -> false + } } diff --git a/core/src/test/java/app/vela/core/data/LowRamModeTest.kt b/core/src/test/java/app/vela/core/data/LowRamModeTest.kt new file mode 100644 index 00000000..02c7911e --- /dev/null +++ b/core/src/test/java/app/vela/core/data/LowRamModeTest.kt @@ -0,0 +1,85 @@ +package app.vela.core.data + +import org.junit.Assert.assertFalse +import org.junit.Assert.assertTrue +import org.junit.Test + +/** + * The low-RAM predicate gates every memory adaptation in the app, and it has already shipped wrong + * twice: an exclusive `1..127` that skipped 128 - the single heap class the target phones use - and + * a zero-means-roomy fallthrough that sent a device we knew nothing about down the memory-hungry + * path. Neither was catchable without a device that reproduced it, which nobody on the dev side has. + * Hence these. + */ +class LowRamModeTest { + + // ---- the off-by-one that started this ---- + + @Test + fun `heap class 128 is low-RAM, the boundary the first version excluded`() { + assertTrue(LowRamMode.classify(isLowRamDevice = false, heapClassMb = 128, totalRamMb = 3000)) + } + + @Test + fun `heap class 127 and below are low-RAM`() { + assertTrue(LowRamMode.classify(isLowRamDevice = false, heapClassMb = 127, totalRamMb = 3000)) + assertTrue(LowRamMode.classify(isLowRamDevice = false, heapClassMb = 96, totalRamMb = 3000)) + assertTrue(LowRamMode.classify(isLowRamDevice = false, heapClassMb = 1, totalRamMb = 3000)) + } + + @Test + fun `heap class above the ceiling is not low-RAM on its own`() { + assertFalse(LowRamMode.classify(isLowRamDevice = false, heapClassMb = 129, totalRamMb = 3000)) + assertFalse(LowRamMode.classify(isLowRamDevice = false, heapClassMb = 192, totalRamMb = 3000)) + } + + // ---- unknown must never read as roomy ---- + + @Test + fun `both probes unreadable is treated as constrained, not roomy`() { + // 0 means "could not read it". Failing to the low-RAM path costs a roomy phone about a + // second on its first mic tap; failing the other way can OOM a phone with no headroom. + assertTrue(LowRamMode.classify(isLowRamDevice = false, heapClassMb = 0, totalRamMb = 0)) + } + + @Test + fun `one unreadable probe does not veto a good reading from the other`() { + // Heap class unknown but 3 GB of RAM: roomy. + assertFalse(LowRamMode.classify(isLowRamDevice = false, heapClassMb = 0, totalRamMb = 3000)) + // RAM unknown but a 256 MB heap class: roomy. + assertFalse(LowRamMode.classify(isLowRamDevice = false, heapClassMb = 256, totalRamMb = 0)) + // RAM unknown and a small heap class: constrained. + assertTrue(LowRamMode.classify(isLowRamDevice = false, heapClassMb = 96, totalRamMb = 0)) + } + + // ---- total RAM, the signal that actually describes the device ---- + + @Test + fun `a nominal 2 GB phone is low-RAM even with a generous heap class`() { + // totalMem reports what the OS can hand out, so a 2 GB phone lands near 1900 MB. + assertTrue(LowRamMode.classify(isLowRamDevice = false, heapClassMb = 192, totalRamMb = 1900)) + } + + @Test + fun `total RAM boundary is inclusive`() { + assertTrue(LowRamMode.classify(isLowRamDevice = false, heapClassMb = 256, totalRamMb = 2048)) + assertFalse(LowRamMode.classify(isLowRamDevice = false, heapClassMb = 256, totalRamMb = 2049)) + } + + // ---- the canonical flag always wins ---- + + @Test + fun `isLowRamDevice alone is enough however roomy the other probes look`() { + assertTrue(LowRamMode.classify(isLowRamDevice = true, heapClassMb = 512, totalRamMb = 8000)) + } + + // ---- the device every measurement in AGENTS.md was taken on ---- + + @Test + fun `the M5 dev phone stays on the normal path`() { + // Measured on device: heapClassMb=256, totalRamMb=2878 (MemTotal 2947424 kB). If this ever + // flips, every memory figure recorded in AGENTS.md was taken on a different code path than + // the one that ships to a roomy phone. + assertFalse(LowRamMode.classify(isLowRamDevice = false, heapClassMb = 256, totalRamMb = 2878)) + } +} diff --git a/docs/FEATURES.md b/docs/FEATURES.md index 8814d10c..df0ca252 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -285,7 +285,7 @@ Status legend: [x] done · [~] partial / in progress · [ ] planned - [x] **Opt-in diagnostics/debug export** (Settings → Diagnostics, off by default) - a local-only event log exportable to JSON via the share sheet, never auto-uploaded, wiped when turned off, in-memory only. - [x] **Crash/ANR/jank capture, all local** - an uncaught-exception handler persists stack traces + breadcrumbs to disk for export, ApplicationExitInfo harvests ANR/native/low-memory kills, a debug ANR watchdog and StrictMode flag stalls and main-thread I/O; captured even with diagnostics off, never auto-sent. - [x] **Memory use cut across the board** (issue #83) - the on-device speech model costs ~267 MB while loaded and used to stay resident for the entire session; it is now dropped after two minutes unused and rebuilt on next use, which reclaims ~101 MB on ANY phone at no visible cost. Every large or native allocation also releases when the system reports memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and gave back nothing when asked. Measured on a 2.9 GB test device: idle memory down 29%, post-trim down 38%, native heap down 57%. Freed memory is also handed back to the operating system rather than kept on the allocator's free lists, which a release alone does not do (a measured ~3 MB more returned per memory warning, at no cost to responsiveness). The two hidden browser views that a search warms up in advance are now released after five idle minutes instead of being held for the whole session: on a 2.9 GB test device that returns about 390 MB to the app and shuts down a separate ~305 MB browser renderer process, with the views rebuilt automatically the next time a place is opened. The hidden view a search warms up in advance is also no longer given its full working size until a place is actually opened, since the warm-up page has nothing to read off it: that cuts about 390 MB while you search and browse, and the gallery is built from exactly the same data as before (verified: the same place returned the same 28 photos before and after). It also drops the scraped page as soon as the photos have been read instead of holding it for two minutes, which cuts a further ~440 MB while you are reading a place, again with the same photos. -- [x] **Low-RAM device adaptation** (issue #83) - additionally, the app detects a memory-constrained phone (`isLowRamDevice` or a heap class of 128 MB or less) and adapts: the on-device speech model is not preloaded at startup, the image cache is capped at 16 MB instead of 48 MB, WebView renderers are not warmed speculatively, and the ambient POI fan-out asks for a smaller result pool per category. Constrained phones still search every category the roomier ones do, so no kind of place (a gym, a bar, a pharmacy, a school, a park) is ever missing from the map because of the device. Independently, every large or native allocation now releases on OS memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and returned nothing when the OS asked. Measured on a 2.9 GB test device: peak memory down 30%, post-trim memory down 38%, native heap down 57%, cold start ~480 ms faster. x86 and x86_64 native libraries are no longer shipped (no target phone can run them), cutting 23 MB of dead weight from every install. +- [x] **Low-RAM device adaptation** (issue #83) - additionally, the app detects a memory-constrained phone (the system low-RAM flag, 2 GB of RAM or less, or a heap class of 128 MB or less, and it assumes constrained when it cannot tell) and adapts: the on-device speech model is not preloaded at startup, the image cache is capped at 16 MB instead of 48 MB, WebView renderers are not warmed speculatively, and the ambient POI fan-out asks for a smaller result pool per category. Constrained phones still search every category the roomier ones do, so no kind of place (a gym, a bar, a pharmacy, a school, a park) is ever missing from the map because of the device. Independently, every large or native allocation now releases on OS memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and returned nothing when the OS asked. Measured on a 2.9 GB test device: peak memory down 30%, post-trim memory down 38%, native heap down 57%, cold start ~480 ms faster. x86 and x86_64 native libraries are no longer shipped (no target phone can run them), cutting 23 MB of dead weight from every install. - [x] **Trip recording + replay** (Settings → "Save my trips", off by default, separate opt-in) - records each drive's GPS trace to a local file replayable on the map at 3x through the real nav pipeline; saved on arrival, listed with Replay/Share/Delete, Share exporting the raw CSV; replay auto-routes to the destination. - [x] **Simulate driving (demo mode)** (Settings → Navigation, off by default) - Start drives any planned route as a synthetic GPS trace through the live-nav loop so nav runs anywhere for demos and screenshots; End stops it; turn off to navigate for real. - [x] **Simulate my location (demo mode)** (off by default) - Vela pretends you're at the map centre so the dot, directions origin and recenter read from there; turn off for real GPS. From d022056a9ddbeae646c87d632c621049a0911e76 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Tue, 21 Jul 2026 12:27:16 -0400 Subject: [PATCH 10/21] Verify the low-RAM path on a production build, and stress it for real Two gaps closed, both methodology, both with numbers that did not exist before. FIRST: the low-RAM branches had never run on a minified build. debug.vela.lowram is BuildConfig.DEBUG-gated, so it is inert on staging and release, and every low-RAM measurement so far came from a debug build with the flag forced. Magisk's resetprop closes that: adb shell su -c "resetprop dalvik.vm.heapgrowthlimit 96m" ActivityManager.staticGetMemoryClass() reads that property per call, so the app sees the new heap class immediately and LowRamMode.classify runs for real (forced=no). Restore it afterwards - it sets the heap growth limit for every app started after. Verifying on staging needs a behavioural probe rather than a log, because Timber.plant(DebugTree) is also DEBUG-gated and MemoryPressure never reaches logcat there. The renderer count works: low-RAM skips the speculative WebView warm, so `ps -A | grep -c sandboxed_process` is 0 after a search on that path and 1 on the normal one. Across three pairs that signal was 0/0/0 against 1/1/1, perfectly separated where PSS was not. Measured that way, the low-RAM path is worth about 148 MB on the production build: 166/169/173 MB against 287/372/295 MB, no crashes, which also shows R8 does not break those branches. SECOND: nothing had ever put the app under real memory pressure. A hog that allocates once measures nothing - this device has 1.6 GB of zram, so the pages are simply compressed and MemAvailable goes UP. The first attempt "applied" 1500 MB and freed 148 MB. Re-touching every page in a loop denies them to the swapper, and then it bites: MemAvailable fell to ~130 MB and lmkd started killing on "direct reclaim and thrashing". What that showed is worth recording, because it cuts against the design this branch inherited. Under real pressure staging received ZERO onTrimMemory callbacks and was killed 20 s in, at oom_score_adj 0, reason "device is not responding". The debug build at a gentler 1.6 GB got EXACTLY ONE level=15, released and purged in 1 ms, and was killed 8 s later. lmkd kills on thrash-driven unresponsiveness before AMS gets round to asking anyone to release anything. So a release that only happens on a trim mostly does not happen. Proactive reclaim - the idle reapers, not warming what will not be used, not holding a scraped document - is what actually protects a constrained phone. That is an argument for the changes on this branch over the trim fan-out alone, and it is recorded here rather than claimed in a commit subject. Also noted: do not push past ~1.6 GB on this hardware. At 1.9 GB the launcher enters a kill loop and the app dies before MemoryPressure.init runs, so the test stops discriminating between good and bad memory behaviour. A continuously rewritten 1.9 GB is a pathological workload, not a small phone. Docs only; no behaviour change. --- AGENTS.md | 36 ++++++++++++++++++++++++++++++++++++ 1 file changed, 36 insertions(+) diff --git a/AGENTS.md b/AGENTS.md index bf5088e2..ca96e615 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -662,6 +662,42 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): phone reads ~1900 MB and a 3 GB phone ~2800 MB. The M5 reads 2878 MB and stays on the normal path, which is what keeps every measurement in this file comparable - there is a test asserting exactly that, so if it ever flips you will be told. + - **`resetprop` exercises REAL low-RAM detection, and is the ONLY way to do it on a release + build.** `debug.vela.lowram` is `BuildConfig.DEBUG`-gated, so it is inert on `staging`/`release` + - the low-RAM branches had therefore never run on a minified build at all. + + adb shell su -c "resetprop dalvik.vm.heapgrowthlimit 96m" # then relaunch + adb shell su -c "resetprop dalvik.vm.heapgrowthlimit 256m" # restore + + `ActivityManager.staticGetMemoryClass()` reads that property per call, so the app sees the new + heap class immediately and `LowRamMode.classify` runs for real (`forced=no`). Magisk's + `resetprop` is what makes a `ro.`-style property writable. RESTORE IT - it changes the heap + growth limit for every app started afterwards. + - **Verify by BEHAVIOUR, not by log, on staging.** `Timber.plant(DebugTree)` is DEBUG-gated, so + `MemoryPressure init` never reaches logcat on a production build. Use the renderer count + instead: low-RAM skips the speculative WebView warm, so + `adb shell ps -A | grep -c sandboxed_process` is 0 after a search on the low-RAM path and 1 on + the normal one. That signal was 0/0/0 versus 1/1/1 across three pairs - perfectly separated, + unlike PSS. + - Measured this way on `staging`, the low-RAM path is worth about **148 MB**: 166/169/173 MB + against 287/372/295 MB. It also proves R8 did not break those branches (0 crashes). + - **Simulating memory pressure: the pages must stay HOT or you measure nothing.** A hog that + allocates once is simply compressed into this device's 1.6 GB of zram, and `MemAvailable` goes + UP - the first attempt at this "applied" 1500 MB and freed 148 MB. Re-touch every page in a loop + to deny them to the swapper. Then it bites: `MemAvailable` fell to ~130 MB and lmkd began + killing on "direct reclaim and thrashing". + - **What that revealed, and it matters for the whole design: trims are not a reliable defence.** + Under real pressure `staging` received **zero** `onTrimMemory` callbacks and was killed 20 s + in (`oom_score_adj 0`, reason "device is not responding"); the debug build at a gentler 1.6 GB + got **exactly one** `level=15`, released and purged in 1 ms, and was killed 8 s later. lmkd + kills on thrash-driven unresponsiveness before AMS gets round to asking anyone to release. + **Proactive reclaim - the idle reapers, not warming what will not be used, not holding a + document after a scrape - is what actually protects a constrained phone.** A release that only + happens on trim mostly does not happen. + - Do not push past ~1.6 GB on this device. At 1.9 GB the launcher enters a kill loop and the app + dies before `MemoryPressure.init` even runs, so the test stops discriminating between good and + bad memory behaviour and only says "the device is broken". A continuously-rewritten 1.9 GB is + a pathological workload, not a small phone. - **Verify the low-RAM path or it ships unverified.** Every dev phone we own reports `lowRam=false heapClassMb=256`, so those branches are dead code locally. Debug builds honour `adb shell setprop debug.vela.lowram true` (then relaunch). NB `false` FORCES the normal path, From 8e710810c4f460fcaaabd65effcb9b7bd4b4eb4b Mon Sep 17 00:00:00 2001 From: alltechdev Date: Tue, 21 Jul 2026 12:57:08 -0400 Subject: [PATCH 11/21] Build the camera index without allocating 400,000 throwaway objects Startup on this app is GC-bound, not compilation-bound. An atrace of a cold start on the staging build attributes seconds to GC phases (CopyingPhase, NativeAlloc concurrent copying GC, MarkingPhase), and forcing full AOT compilation made cold start WORSE, 828 ms against 775 ms - so a baseline profile would buy nothing and the thing worth cutting is allocation. FlockCameras.loadFrom was the largest single allocator at startup. For the bundled dataset (127,770 rows as shipped today) it created roughly 400,000 objects that were garbage within milliseconds: - 255,540 boxed java.lang.Double, from two ArrayList accumulators that were copied into DoubleArrays and discarded - one boxed Integer per camera, 127,770 of them, from .add(i) into MutableList - per occupied cell an ArrayList, its Object[] backing, a boxed Long key and a HashMap.Node - 14,273 cells, so ~57,000 more The coordinates now accumulate into primitive DoubleArrays that double on demand, and the 0.1 degree bucket index is a flat CSR triple - sorted unique LongArray of cell keys, IntArray of start offsets, IntArray of row indices - built with a primitive sort and two counting passes. Five arrays instead of ~200,000 objects. Lookup is a binary search over ~14,000 keys, and the two call sites move from `grid[key(r, c)]?.let { for (i in bucket) ... }` to an inline forEachInCell, which also drops the iterator and the unboxing per step. Proven equivalent on the real dataset rather than argued: a temporary check rebuilt the old HashMap index alongside the new one during a real load on device and compared every bucket. rows=127770 cells=14273 oldCells=14273 mismatches=0. The check was removed again after it passed; this is the only kind of evidence worth having for an index rewrite, since the data is the thing that finds the edge cases. On the GC metric it was written for, mean GC time across a cold start fell from 3408 ms to 2340 ms (n=3 per arm). Stated honestly, that is directional, not conclusive: the arms overlap (2355-4644 before, 1641-3144 after) and three runs cannot separate them. Cold start wall time is worse still as a metric here - it swings 739-1552 ms on this device - which is why GC slices were measured instead. The allocation reduction itself is not in doubt; the size of its effect is. Behaviour is unchanged by construction: same rows, same cells, same buckets, same camera list. The Flock layer is ON by default (a deliberate 2026-07-13 call), so deferring this work instead was not an option - it had to get cheaper, not later. Verified: :app:detekt 0 smells, :core:test green, audit_deadcode.sh PASS, standardStaging builds, installs and runs, map renders, 0 crashes. --- .../main/java/app/vela/data/FlockCameras.kt | 91 +++++++++++++++---- 1 file changed, 73 insertions(+), 18 deletions(-) diff --git a/app/src/main/java/app/vela/data/FlockCameras.kt b/app/src/main/java/app/vela/data/FlockCameras.kt index 1b081343..8ed0d59f 100644 --- a/app/src/main/java/app/vela/data/FlockCameras.kt +++ b/app/src/main/java/app/vela/data/FlockCameras.kt @@ -38,7 +38,24 @@ object FlockCameras { private var lat = DoubleArray(0) private var lng = DoubleArray(0) private var op = arrayOf() - private val grid = HashMap>() + /** + * The 0.1 deg bucket index, as a flat CSR (compressed sparse row) triple rather than a + * `HashMap>`. + * + * [cellKeys] is sorted and unique; the rows in cell `k` are + * `cellRows[cellStart[k] until cellStart[k + 1]]`. Lookup is a binary search on [cellKeys]. + * + * Why: the map version cost roughly 166,000 objects to build - one boxed `Integer` per camera + * (124,406 of them), plus an `ArrayList` + its `Object[]` + a boxed `Long` key + a `HashMap.Node` + * per occupied cell (13,965 of those) - and every one of them was garbage the moment the parse + * finished. This is five arrays. Startup on this app is GC-bound, not compilation-bound (forcing + * full AOT made cold start WORSE, 828 ms against 775 ms), so allocation count at startup is the + * thing worth cutting. Lookups are unchanged in behaviour and no slower in practice: the viewport + * scan touches on the order of 100 cells, and a binary search over ~14,000 keys is ~14 compares. + */ + private var cellKeys = LongArray(0) + private var cellStart = IntArray(1) + private var cellRows = IntArray(0) val isLoaded: Boolean get() = loaded val size: Int get() = lat.size @@ -46,6 +63,19 @@ object FlockCameras { private fun key(row: Long, col: Long): Long = (row shl 32) xor (col and 0xffffffffL) private fun rowOf(v: Double): Long = Math.floor(v / CELL).toLong() + /** + * Run [body] for every camera row in the cell at ([row], [col]). No-op when the cell is empty. + * Replaces `grid[key(r, c)]?.let { ... }`, and allocates nothing - notably no iterator, which + * the old `for (i in bucket)` over a `MutableList` created (and unboxed on every step). + */ + private inline fun forEachInCell(row: Long, col: Long, body: (Int) -> Unit) { + val k = cellKeys.binarySearch(key(row, col)) + if (k < 0) return + var p = cellStart[k] + val end = cellStart[k + 1] + while (p < end) { body(cellRows[p]); p++ } + } + private fun dir(context: Context) = File(context.filesDir, "flock").apply { mkdirs() } private fun downloadedBin(context: Context) = File(dir(context), "cameras.bin") private fun downloadedVer(context: Context) = File(dir(context), "version.txt") @@ -72,9 +102,13 @@ object FlockCameras { /** Build the arrays + index from a gzipped-TSV stream and publish them (never leaves `loaded` false once set). */ private fun loadFrom(raw: InputStream) { - val las = ArrayList(130_000) - val los = ArrayList(130_000) - val ops = ArrayList(130_000) + // Primitive growable arrays, not ArrayList: the list version boxed a java.lang.Double + // per coordinate, 248,812 of them for the bundled 124,406-row dataset, all garbage the moment + // toDoubleArray() copied them out. + var las = DoubleArray(130_000) + var los = DoubleArray(130_000) + var ops = arrayOfNulls(130_000) + var n = 0 val intern = HashMap() // operator column is highly repetitive - intern it raw.use { r -> GZIPInputStream(r).bufferedReader().useLines { lines -> @@ -84,15 +118,40 @@ object FlockCameras { val la = line.substring(0, t1).toDoubleOrNull() ?: continue val lo = line.substring(t1 + 1, t2).toDoubleOrNull() ?: continue val o = line.substring(t2 + 1) - las.add(la); los.add(lo); ops.add(intern.getOrPut(o) { o }) + if (n == las.size) { // dataset outgrew the guess - double, same as ArrayList did + las = las.copyOf(n * 2); los = los.copyOf(n * 2); ops = ops.copyOf(n * 2) + } + las[n] = la; los[n] = lo; ops[n] = intern.getOrPut(o) { o } + n++ } } } - val g = HashMap>() - for (i in las.indices) g.getOrPut(key(rowOf(las[i]), rowOf(los[i]))) { ArrayList() }.add(i) + + // CSR index, built with sorts and counting rather than a map of lists. Three passes over + // primitives and no per-row object at all; see [cellKeys]. + val keys = LongArray(n) { key(rowOf(las[it]), rowOf(los[it])) } + val sorted = keys.copyOf() + sorted.sort() + var uniq = 0 + for (i in 0 until n) if (i == 0 || sorted[i] != sorted[i - 1]) uniq++ + val ck = LongArray(uniq) + var u = 0 + for (i in 0 until n) if (i == 0 || sorted[i] != sorted[i - 1]) { ck[u] = sorted[i]; u++ } + // Count per cell, then prefix-sum into start offsets. + val start = IntArray(uniq + 1) + for (i in 0 until n) start[ck.binarySearch(keys[i]) + 1]++ + for (k in 1..uniq) start[k] += start[k - 1] + // Scatter row indices into their cell's slot. `fill` walks a copy of the offsets so + // `start` stays the published boundary array. + val fill = start.copyOf() + val rows = IntArray(n) + for (i in 0 until n) { val k = ck.binarySearch(keys[i]); rows[fill[k]] = i; fill[k]++ } + // Publish (a bad/partial parse threw before here, so we never swap in a half-built set). - lat = las.toDoubleArray(); lng = los.toDoubleArray(); op = ops.toTypedArray() - grid.clear(); grid.putAll(g) + lat = las.copyOf(n); lng = los.copyOf(n) + @Suppress("UNCHECKED_CAST") + op = (ops.copyOf(n) as Array) + cellKeys = ck; cellStart = start; cellRows = rows loaded = true } @@ -140,10 +199,8 @@ object FlockCameras { while (r <= r1) { var c = c0 while (c <= c1) { - grid[key(r, c)]?.let { bucket -> - for (i in bucket) { - if (lat[i] in south..north && lng[i] in west..east) out.add(AlprCamera(LatLng(lat[i], lng[i]), op[i])) - } + forEachInCell(r, c) { i -> + if (lat[i] in south..north && lng[i] in west..east) out.add(AlprCamera(LatLng(lat[i], lng[i]), op[i])) } c++ } @@ -163,11 +220,9 @@ object FlockCameras { while (r <= r1) { var c = c0 while (c <= c1) { - grid[key(r, c)]?.let { bucket -> - for (i in bucket) { - val p = LatLng(lat[i], lng[i]) - if (nearPolyline(p, polyline, meters)) out.add(AlprCamera(p, op[i])) - } + forEachInCell(r, c) { i -> + val p = LatLng(lat[i], lng[i]) + if (nearPolyline(p, polyline, meters)) out.add(AlprCamera(p, op[i])) } c++ } From f1d21ae7a5a65790cc18b7f8bb98bfb75a8c17c3 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Tue, 21 Jul 2026 13:01:08 -0400 Subject: [PATCH 12/21] Record what startup profiling actually found, including the biggest unfixed item Four measurements worth not repeating, and one large item this branch does NOT fix. MapScreen is too big for ART to compile. On the shipping build ART logs "Method exceeds compiler instruction limit: 19621 in void i2.r1.f(...)", which the R8 mapping resolves to MapScreenKt.MapScreen - so the composable that runs on every recomposition of the main screen is never compiled and runs interpreted. Its body spans roughly lines 192-2172. This is the largest known performance item in the app and it is still open. A trial extraction of the biggest block (lines 1828-2048, 221 lines) was done and reverted, and the recipe is recorded so the next attempt is not exploratory: the extracted function needs a BoxScope receiver, it captures 23 named values, and six of those are `var ... by remember` that the block WRITES - pass the MutableState and re-delegate with `by` so the body stays byte-identical. Passing them by value is a compile error rather than a silent break, which is what makes the refactor tractable at all. It should be done one block at a time, re-checking the logcat instruction count after each. Startup is GC-bound, not compilation-bound. Forcing full AOT made cold start WORSE, 828 ms against 775 ms, which is an upper bound on anything a baseline profile could buy - so that idea is closed rather than pending. An atrace instead attributes seconds to GC phases, which is what motivated the FlockCameras index rewrite in the previous commit. Measuring startup needs the dexopt state controlled or it measures nothing: `adb install -r` resets it, and comparing a fresh-install arm against a warmed baseline once "showed" that REMOVING work made startup slower. Cold start swings 739-1552 ms on this device even when matched, so anything smaller than a few hundred ms needs a lower-variance metric. And dumpsys gfxinfo does not measure this app's map at all - MapLibre renders through its own GL context, so HWUI stats cover only the Compose chrome. A D-pad drive produced 60 frames at 0% jank and a swipe drive 11 frames; neither could have detected a regression in either direction. Docs only; no behaviour change. --- AGENTS.md | 43 +++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 43 insertions(+) diff --git a/AGENTS.md b/AGENTS.md index ca96e615..f2c2b405 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -905,6 +905,49 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): ~11 MB instead of ~111 MB, and a memory benchmark silently measures the model-absent case. Check `secondary` is actually ~111 MB before believing any ASR memory number. +## Performance: what has been measured + +- **`MapScreen` is too big for ART to COMPILE, so the main screen runs interpreted.** On the + shipping build ART logs `Method exceeds compiler instruction limit: 19621 in void + i2.r1.f(i2.E3, z3.a, Z.p, int)`, which the R8 mapping resolves to + `MapScreenKt.MapScreen(MapViewModel, Function0, Composer, int)` (`MapScreenKt -> i2.r1`, + `MapScreen -> f`). ART's optimizing compiler skips methods over ~10,000 dex instructions, so the + composable that runs on every recomposition of the main screen is never compiled. Verify with: + + adb logcat -d | grep -i "exceeds compiler instruction limit" + + The body spans roughly lines 192-2172 of a 3,479-line file. **This is the largest known + performance item in the app and it is NOT fixed.** + - The fix is to extract the big `if` blocks under the root `Box` into private composables. A + trial extraction of the largest (lines 1828-2048, the idle-map overlay block, 221 lines) was + done and reverted, and it establishes the recipe: the function needs a **`BoxScope` receiver** + (the block uses `Modifier.align`), and it captures **23** names - `chromeLift, context, + darkTheme, driveFollowing, followMe, layersOpen, metersPerPixel, parkedCarLabel, + parkingClearedMsg, parkingMovedMsg, parkingNoFixMsg, parkingSet, parkingTapAction, resultsShown, + searchOpen, show, showParkingHistory, showParkingMenu, softkeyBarShown, speedOverlayArmed, + state, vm`. + - Six of those are `var ... by remember { mutableStateOf(...) }` (`followMe`, `layersOpen`, + `metersPerPixel`, `showParkingHistory`, `showParkingMenu`, `speedOverlayArmed`) and the block + WRITES them. Pass the `MutableState` and re-delegate at the top of the extracted function + (`var showParkingMenu by showParkingMenuState`) so the 221-line body stays byte-identical. + Passing them by value is a compile error, not a silent break - the compiler is the safety net + here, which is what makes this refactor tractable. + - Do it ONE block at a time, rebuilding and re-checking the logcat number after each, and stop + when the message disappears. Screenshot the map after each step. +- **Startup is GC-bound, not compilation-bound. Do not reach for a baseline profile.** Forcing full + AOT (`cmd package compile -m speed -f`) made cold start WORSE - 828 ms against 775 ms - which is + an upper bound on anything a profile could buy. An atrace of a cold start instead attributes + seconds to GC (`CopyingPhase`, `NativeAlloc concurrent copying GC`, `MarkingPhase`), so allocation + count at startup is the thing worth cutting. That is what motivated the `FlockCameras` CSR index. +- **Measuring startup: control the dexopt state or measure nothing.** `adb install -r` resets it, so + runs straight after an install are unprofiled and slower. Comparing a fresh-install arm against a + warmed baseline once "showed" that REMOVING work made startup slower. Cold start on this device + swings 739-1552 ms even matched, so prefer a lower-variance metric (atrace GC slices) for anything + smaller than a few hundred ms. +- **`dumpsys gfxinfo` does not measure this app's map.** MapLibre renders through its own GL context, + so HWUI frame stats cover only the Compose chrome - a D-pad drive produced 60 frames at 0% jank + and a swipe drive 11 frames, neither of which could have detected a regression. + ## Layout - `:core` is the UI-agnostic "extractor" (NewPipeExtractor pattern). `:app` is From bec5279e4a108666dc404eb8ae77d23af58a5826 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Tue, 21 Jul 2026 13:06:59 -0400 Subject: [PATCH 13/21] Record that the obvious MapScreen extraction makes it worse The previous note said MapScreen exceeds ART's compiler limit and described how to extract a block out of it. Following that advice naively does not work, and this is measured rather than reasoned. The largest block (lines 1828-2048, the idle-map overlay cluster, 221 lines) was fully extracted into BoxScope.MapIdleOverlays with all 21 captures passed and the six `var ... by remember` states passed as MutableState and re-delegated so the body stayed byte-identical. It compiled, installed, ran and did not crash. And MapScreen went from 19,621 to 20,351 instructions - the wrong direction. The reason is Compose's calling convention: it emits $changed/$changed1 bitmask plumbing per parameter at the call site, and for 21 parameters that costs more than a 221-line body removes. The change was reverted; the tree is back to 19,621 and green. So the rule is the opposite of the intuitive one: extract blocks with FEW captures, not the biggest blocks. Count the captures first - comment the block out, compile, and read the unresolved references, which is the exact list for one build. A 100-line block taking 4 parameters beats a 220-line block taking 21. Recomputing composable-local values inside the extracted function (LocalContext.current, stringResource, isAppInDarkTheme, VelaSoftkeys.isActive) instead of passing them drops the count further at no behavioural cost. Docs only; no behaviour change. MapScreen is still uncompiled and still the largest known performance item. --- AGENTS.md | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/AGENTS.md b/AGENTS.md index f2c2b405..13bc0ab7 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -934,6 +934,18 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): here, which is what makes this refactor tractable. - Do it ONE block at a time, rebuilding and re-checking the logcat number after each, and stop when the message disappears. Screenshot the map after each step. + - **A naive extraction makes it WORSE, and this was measured, not guessed.** The 221-line block + above was fully extracted into `BoxScope.MapIdleOverlays` with all 21 captures passed and the + six `MutableState`s re-delegated. It compiled, ran and did not crash - and MapScreen went + **19,621 -> 20,351 instructions**. Compose emits `$changed`/`$changed1` bitmask plumbing per + parameter at the call site, and for a capture set that size it costs more than the body removes. + The change was reverted. + - So the rule is: **extract blocks with FEW captures, not the biggest blocks.** Before extracting, + count the captures (comment out the block, compile, read the unresolved references - that is + the exact list, and it takes one build). A 100-line block taking 4 parameters will beat a + 220-line block taking 21. Recomputing composable-local values inside the extracted function + (`LocalContext.current`, `stringResource`, `isAppInDarkTheme()`, `VelaSoftkeys.isActive()`) + rather than passing them also drops the count without changing behaviour. - **Startup is GC-bound, not compilation-bound. Do not reach for a baseline profile.** Forcing full AOT (`cmd package compile -m speed -f`) made cold start WORSE - 828 ms against 775 ms - which is an upper bound on anything a profile could buy. An atrace of a cold start instead attributes From fc62f01b3b46bb3c70d1ef029b083221d8615443 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Tue, 21 Jul 2026 13:12:42 -0400 Subject: [PATCH 14/21] Close off composable extraction as a way to fix MapScreen The previous note said extraction makes MapScreen worse and blamed Compose's per-parameter $changed plumbing, concluding that blocks with FEW captures should be extracted instead. That follow-up hypothesis was tested and is also wrong. block extracted lines caps lines/cap MapScreen after (baseline) - - - 19,621 1828-2048 idle overlays 221 23 9.6 20,351 (+730) 991-1104 dpad overlay 114 9 12.7 21,638 (+2,017) The second block was picked precisely because its lines-per-capture ratio was far better, which the parameter-plumbing theory predicted would win. It came out worse than the first, and worse in absolute terms. Both compiled, installed and ran with no crashes, so this is purely a code-size result; both were reverted and the tree is back at 19,621. Whatever dominates the instruction count, adding a composable call layer costs more than the body it removes. Extraction is therefore closed as an approach: MapScreen cannot be brought under the ~10,000 limit by pulling blocks out of it, and a third attempt at the same idea is not worth anyone's build time. The mechanics recorded earlier (BoxScope receiver, passing MutableState and re-delegating with `by` to keep the body byte-identical) are correct and worth keeping - it is the strategy they served that does not hold. MapScreen remains uncompiled and remains the largest known performance item. A different hypothesis is needed, and it should be measured in one build before any refactoring work is done. Docs only; no behaviour change. --- AGENTS.md | 20 ++++++++++++++++++++ 1 file changed, 20 insertions(+) diff --git a/AGENTS.md b/AGENTS.md index 13bc0ab7..95ba9fa2 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -934,6 +934,26 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): here, which is what makes this refactor tractable. - Do it ONE block at a time, rebuilding and re-checking the logcat number after each, and stop when the message disappears. Screenshot the map after each step. + - **EXTRACTION DOES NOT WORK. Two controlled experiments, both worse. Do not try a third.** + Extracting a block into a private composable consistently INCREASES MapScreen's instruction + count, and the effect is not explained by how many values the block captures: + + | block extracted | lines | captures | lines/capture | MapScreen after | + |---|---|---|---|---| + | (baseline) | - | - | - | **19,621** | + | 1828-2048 idle overlays | 221 | 23 | 9.6 | 20,351 (+730) | + | 991-1104 dpad overlay | 114 | 9 | 12.7 | **21,638 (+2,017)** | + + The second was chosen precisely because it had a far better lines-per-capture ratio, on the + theory that Compose's per-parameter `$changed` plumbing was the cost. It came out WORSE than the + first. Both compiled, installed and ran without crashing, so this is a code-size result, not a + correctness one; both were reverted. Whatever dominates the count, adding a composable call + layer costs more than the body it removes. **Shrinking MapScreen under the ~10,000 limit is not + reachable by pulling blocks out of it**, and anyone trying should have a different hypothesis + and measure it in one build before doing the work. + - The stale earlier advice below is kept only for the mechanics it records (BoxScope receiver, + passing `MutableState` and re-delegating with `by` so the body stays byte-identical). Those + techniques are correct; the strategy they serve is not. - **A naive extraction makes it WORSE, and this was measured, not guessed.** The 221-line block above was fully extracted into `BoxScope.MapIdleOverlays` with all 21 captures passed and the six `MutableState`s re-delegated. It compiled, ran and did not crash - and MapScreen went From b596ededf541d644d529ed00cb0d07ab931ffe40 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Tue, 21 Jul 2026 13:16:34 -0400 Subject: [PATCH 15/21] The UI is not the bottleneck: 9.0 ms per frame, against a 16.7 ms budget MapScreen exceeding ART's compiler limit looked like the largest performance item in the app, and two extraction attempts were made and reverted chasing it. Nobody had checked whether it actually costs frames. It does not. An atrace across a cold start plus a D-pad drive on the staging build: Choreographer#doFrame 3,860 ms over 430 frames = 9.0 ms/frame measure/layout/draw 1,524 ms = 3.5 ms/frame GC 2,165 ms inflate 16 ms Against a 16.7 ms budget at 60 Hz there is no frame-budget problem on this device. A method being uncompiled only matters if it runs hot, and at 9 ms/frame this one is not hurting. That explains why both extractions changed the instruction count and nothing a user could feel: they were aimed at something that was not costing anything. This measurement should have been taken first, and the note now says so. It also redirects the effort: GC is the largest remaining cost in the trace, which is what the FlockCameras index rewrite targeted, and allocation is where further work belongs rather than composable restructuring. MapScreen stays recorded as the largest code-size anomaly, but is explicitly demoted from "largest performance item" to "do not spend effort here without first showing it costs frames". Docs only; no behaviour change. --- AGENTS.md | 14 ++++++++++++-- 1 file changed, 12 insertions(+), 2 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 95ba9fa2..2e40d053 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -907,6 +907,15 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): ## Performance: what has been measured +- **The UI is NOT the bottleneck - measure before optimising it.** An atrace across cold start plus + a D-pad drive on the staging build: `Choreographer#doFrame` totals 3,860 ms over **430 frames = + 9.0 ms per frame**, against a 16.7 ms budget at 60 Hz. measure/layout/draw is 3.5 ms/frame and + inflate is 16 ms in total. **There is no frame-budget problem on this device.** GC, at 2,165 ms, + is the largest remaining cost, which is why allocation work (see the `FlockCameras` index) is the + productive direction and composable restructuring is not. + - This measurement should have come FIRST. Two `MapScreen` extractions were attempted and reverted + before anyone checked whether the uncompiled composable was actually costing frames. It is not. + A method being uncompiled only matters if it runs hot, and at 9 ms/frame this one does not hurt. - **`MapScreen` is too big for ART to COMPILE, so the main screen runs interpreted.** On the shipping build ART logs `Method exceeds compiler instruction limit: 19621 in void i2.r1.f(i2.E3, z3.a, Z.p, int)`, which the R8 mapping resolves to @@ -916,8 +925,9 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): adb logcat -d | grep -i "exceeds compiler instruction limit" - The body spans roughly lines 192-2172 of a 3,479-line file. **This is the largest known - performance item in the app and it is NOT fixed.** + The body spans roughly lines 192-2172 of a 3,479-line file. It is the largest known code-size + anomaly, but per the frame measurement above it is **not** a demonstrated performance problem - + do not spend effort here without first showing it costs frames. - The fix is to extract the big `if` blocks under the root `Box` into private composables. A trial extraction of the largest (lines 1828-2048, the idle-map overlay block, 221 lines) was done and reverted, and it establishes the recipe: the function needs a **`BoxScope` receiver** From 6ab1c97c149b2fa1fbea78952f4011ae06715cc4 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Tue, 21 Jul 2026 13:20:14 -0400 Subject: [PATCH 16/21] Point performance work at content latency, which is 28.8 s, not frames, which are fine This session measured frame time, cold start, GC and instruction counts, and the honest summary of all of it is that the app is inside its frame budget with room to spare: 9.0 ms per frame against 16.7 ms. Two MapScreen refactors were attempted and reverted chasing an instruction count that turned out not to cost frames at all. The axis that was never measured until now is the one a user actually experiences. On staging, tapping a place and waiting for a complete gallery is 28.8 seconds (35 photos). That is not a stall - the photo scraper polls up to 58 ticks at 500 ms behind a 40 s timeout, and reviews sit behind 45 s - it is the designed shape of scraping a rendered Google page. But it is what "the whole app is a bit slow" almost certainly means, and it dwarfs anything measurable in frames or startup. So the note now says to start there: how long until the user sees the thing they asked for - search results, ambient pins, the gallery, reviews - rather than gfxinfo or cold start, neither of which showed a problem worth fixing. It also records the clearest unpulled lead. nearbyPlaces fans 15 category terms out 4-at-a-time and ends with awaitAll().flatten(), so every ambient pin appears at once after the slowest term, about four network waves in. Streaming each term as it lands would put the first pins on screen roughly 4x sooner for identical total work. That changes what the user sees while loading, so it wants a before/after measurement rather than a drive-by edit. Docs only; no behaviour change. --- AGENTS.md | 14 ++++++++++++++ 1 file changed, 14 insertions(+) diff --git a/AGENTS.md b/AGENTS.md index 2e40d053..c2b257ab 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -907,6 +907,20 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): ## Performance: what has been measured +- **The slowness users feel is CONTENT LATENCY, not frames. Measure that axis first.** + Device-measured on staging: **tapping a place to a complete gallery is 28.8 s** (35 photos). The + scrapers are built around 40 s and 45 s timeouts with a 7 s page-load allowance, and the photo + scraper's own JS polls up to 58 ticks at 500 ms, so ~30 s is the designed shape, not a stall. + Frame time on the same device is 9.0 ms against a 16.7 ms budget - there is nothing to win there. + A user reporting "the whole app is a bit slow" is far more likely describing this than jank. + - So performance work on Vela should start with: how long until the user sees the thing they asked + for? Search results, ambient POI pins, the gallery, reviews. Not `gfxinfo`, not cold start. + - A concrete lead nobody has pulled: `GoogleMapsDataSource.nearbyPlaces` fans out 15 category + terms 4-at-a-time and finishes with `awaitAll().flatten()`, so **every ambient pin appears at + once after the slowest term**, roughly four network waves in. Streaming each term's results as + they land would put first pins on screen ~4x sooner for the same total work. That is a real + user-visible latency change though, so it needs a before/after and a careful look, not a + drive-by edit. - **The UI is NOT the bottleneck - measure before optimising it.** An atrace across cold start plus a D-pad drive on the staging build: `Choreographer#doFrame` totals 3,860 ms over **430 frames = 9.0 ms per frame**, against a 16.7 ms budget at 60 Hz. measure/layout/draw is 3.5 ms/frame and From 443619a6ca409fb53a03fa86238861d3db260ed0 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Tue, 21 Jul 2026 13:21:41 -0400 Subject: [PATCH 17/21] Record why streaming the ambient fan-out is not a drive-by edit The previous note flagged nearbyPlaces' awaitAll as the clearest latency lead: 15 terms 4-at-a-time, every ambient pin appearing at once after the slowest one, so streaming each term would show first pins ~4x sooner. That is still the right lead, but attempting it quickly would reintroduce a bug this codebase already fixed, and the note now says so. nearbyPlaces post-processes the whole merged pool with the SLIM-FLAVOR HEAL. For the first ~3 s of a session Google serves per-place blocks with the review count absent, which zeroes ambientProminence and - quoting the existing comment - "silently broke everything keyed on it: prominence ranking, dot sizing, label tiers - all flat". The heal detects that flavour across the pool and refetches. Painting each term as it lands would put pins on screen before the heal can run, which is precisely that flat-ranking regression. There is also no onPartial on MapDataSource.nearbyPlaces today - photos and reviews have one, ambient does not - so the interface, the ViewModel call site at MapViewModel.kt:3659, the heal, rankAmbientPlaces and the take-N cap have to be designed together rather than patched at the fan-out. One thing does already hold: collision priority is stable across uploads because it is keyed on prominence rather than list index (upstream c35eea33), so repeated partial uploads will not reshuffle icon placement. That removes one of the two obvious hazards, leaving the heal as the real one. Docs only; no behaviour change. --- AGENTS.md | 15 ++++++++++++--- 1 file changed, 12 insertions(+), 3 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index c2b257ab..f69baaff 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -918,9 +918,18 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): - A concrete lead nobody has pulled: `GoogleMapsDataSource.nearbyPlaces` fans out 15 category terms 4-at-a-time and finishes with `awaitAll().flatten()`, so **every ambient pin appears at once after the slowest term**, roughly four network waves in. Streaming each term's results as - they land would put first pins on screen ~4x sooner for the same total work. That is a real - user-visible latency change though, so it needs a before/after and a careful look, not a - drive-by edit. + they land would put first pins on screen ~4x sooner for the same total work. + - **But it is NOT a drive-by edit, and here is the trap.** `nearbyPlaces` post-processes the whole + fan-out with the SLIM-FLAVOR HEAL: for the first ~3 s of a session Google serves per-place blocks + with the review count ABSENT, which zeroes `ambientProminence` and, in the code's own words, + "silently broke everything keyed on it: prominence ranking, dot sizing, label tiers - all flat". + The heal detects that flavour across the merged pool and refetches. Painting each term as it + lands would put pins on screen BEFORE the heal can run, i.e. exactly the flat-ranking bug that + was already fixed once. There is no `onPartial` on `MapDataSource.nearbyPlaces` today (photos and + reviews have one; ambient does not), so the interface, the ViewModel call site at + MapViewModel.kt:3659, the heal, `rankAmbientPlaces` and the take-N cap all have to be worked out + together. Collision priority is at least already stable across uploads (prominence, not list + index - upstream c35eea33), so repeated uploads will not reshuffle placement. - **The UI is NOT the bottleneck - measure before optimising it.** An atrace across cold start plus a D-pad drive on the staging build: `Choreographer#doFrame` totals 3,860 ms over **430 frames = 9.0 ms per frame**, against a 16.7 ms budget at 60 Hz. measure/layout/draw is 3.5 ms/frame and From 452b698e11c9c8fbe82db093c8afe5c3238e70a1 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Thu, 23 Jul 2026 19:42:31 -0400 Subject: [PATCH 18/21] Fix all 12 review findings, fix the armv7 model-load SIGBUS (issue #95), add a low-memory ASR engine The review round on this PR surfaced 12 verified findings; all are fixed: - WhisperRecognizer: the engine-switch rebuild now retires a superseded recognizer while leases are outstanding instead of freeing it under a running decode; lease acquisition is cancellation-safe (a leaked lease permanently disabled every release path); REAP_IDLE_MS is RAM-scaled (120 s low-RAM / 600 s roomy) so roomy phones keep instant mic taps. - PiperSynth: the ASR two-strike native-crash quarantine now wraps the TTS load per voice (issue #95's crash-loop class), and the trim path releases with interrupt = false so a speaking nav prompt is never cut off mid-word. - Web fetchers: all five now decline a merely-severe trim reap while a fetch is in flight and drain pending deferreds on a forced (critical) teardown - no more 20-45 s mutex-holding hangs or emptied galleries; reap bookkeeping is main-thread-confined everywhere. - MapViewModel.deleteAsrEngine runs its lock wait + native free + 154 MB recursive delete on Dispatchers.IO with an optimistic row hide. - FlockCameras publishes the CSR index as one immutable generation behind a single volatile reference - the refresh hot-swap can no longer tear against a viewport scan (AIOOBE). Unit-tested. - GoogleMapsDataSource: the low-RAM !7i30 pool halving is reverted - it silently removed POI ranks 31-60 with no basemap fallback. armv7 (issue #95): the vendored sherpa-onnx AAR moves 1.13.3 -> 1.13.4, whose bundled onnxruntime (1.27.0, from 1.24.3) fixes the BUS_ADRALN unaligned-read SIGBUS that crashed every model load on 32-bit ARM. Device-verified both directions on an M5 forced to --abi armeabi-v7a: TTS (the exact #95 voice) and ASR load, run, release and survive trims in a 32-bit process. CI fetches the new AAR; do-not-downgrade notes in the build file and AGENTS.md. New: AsrEngine.ZIPFORMER_SMALL (28 MB English transducer) as the low-memory candidate for feature phones - NeMo Conformer CTC was tried and rejected at ~760 MB-1.2 GB resident. Zipformer emits ALL-CAPS spoken-form text, so SpeechText.cleanSearchTranscript now lowercases shouting and runs spokenNumbersToDigits, unit-tested inverse text normalization ("ONE TWENTY THREE MAIN STREET" -> "123 main street", ordinal streets, non-English passthrough). detekt 0 findings; 49 unit tests green (14 new). --- .github/workflows/ci.yml | 10 +- AGENTS.md | 88 +++++++++--- app/build.gradle.kts | 12 +- .../main/java/app/vela/data/FlockCameras.kt | 99 ++++++++------ .../main/java/app/vela/ui/map/MapViewModel.kt | 22 ++- .../ui/settings/sections/SearchSettings.kt | 1 + app/src/main/java/app/vela/voice/AsrEngine.kt | 32 ++++- .../main/java/app/vela/voice/PiperSynth.kt | 57 +++++++- .../java/app/vela/voice/WhisperRecognizer.kt | 87 ++++++++++-- .../java/app/vela/web/WebDirectionsFetcher.kt | 36 ++++- .../main/java/app/vela/web/WebPhotoFetcher.kt | 25 +++- .../app/vela/web/WebPopularTimesFetcher.kt | 22 ++- .../java/app/vela/web/WebReviewsFetcher.kt | 36 ++++- .../app/vela/web/WebStopDeparturesFetcher.kt | 36 ++++- app/src/main/res/values/strings.xml | 1 + .../java/app/vela/data/FlockCamerasTest.kt | 108 +++++++++++++++ .../core/data/google/GoogleMapsDataSource.kt | 20 ++- .../java/app/vela/core/voice/SpeechText.kt | 129 ++++++++++++++++++ .../app/vela/core/voice/SpeechTextTest.kt | 51 +++++++ 19 files changed, 737 insertions(+), 135 deletions(-) create mode 100644 app/src/test/java/app/vela/data/FlockCamerasTest.kt diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index b2dd0a2f..f037dbe6 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -136,15 +136,17 @@ jobs: run: | mkdir -p app/libs # Hosted on this repo's fixed-tag `tts-runtime` release (a one-time manual upload). - curl -fSL -o app/libs/sherpa-onnx-1.13.3.aar \ - "https://github.com/${{ github.repository }}/releases/download/tts-runtime/sherpa-onnx-1.13.3.aar" - test "$(stat -c%s app/libs/sherpa-onnx-1.13.3.aar)" -gt 50000000 + # 1.13.4 (onnxruntime 1.27.0) is REQUIRED, not preferred: its bundled runtime fixes the + # armv7 unaligned-read SIGBUS that crashed every model load on 32-bit phones (issue #95). + curl -fSL -o app/libs/sherpa-onnx-1.13.4.aar \ + "https://github.com/${{ github.repository }}/releases/download/tts-runtime/sherpa-onnx-1.13.4.aar" + test "$(stat -c%s app/libs/sherpa-onnx-1.13.4.aar)" -gt 50000000 # The hosted AAR must actually CARRY 32-bit ARM, not just be big enough. If it is ever # replaced with an arm64-only build, the v7a strip fix (#81) silently reverts: the APK # still builds, still installs on a TCL Flip 2, and voice search + Vela voice are dead # again with an UnsatisfiedLinkError nobody sees until a tester reports it. Fail here # instead - this is the check whose absence let the original bug ship. - unzip -l app/libs/sherpa-onnx-1.13.3.aar | grep -q "jni/armeabi-v7a/libsherpa-onnx-jni.so" \ + unzip -l app/libs/sherpa-onnx-1.13.4.aar | grep -q "jni/armeabi-v7a/libsherpa-onnx-jni.so" \ || { echo "::error::tts-runtime AAR has no armeabi-v7a sherpa-onnx - 32-bit phones would lose voice search + neural TTS"; exit 1; } - name: App unit tests diff --git a/AGENTS.md b/AGENTS.md index 33b7284b..320897da 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -835,9 +835,23 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): - **The ASR model is the single largest reclaimable allocation: ~267 MB PSS** (~101 MB of weights in `scudo:secondary` plus ~146 MB of onnxruntime arena in `scudo:primary`). It releases on a severe trim, is NOT warmed at startup on a low-RAM device, and - on EVERY device - is dropped - after `REAP_IDLE_MS` (120 s) unused and rebuilt on next use. `WhisperRecognizer.release()` - declines while a listen is in flight - freeing the native recognizer under a running decode is a - use-after-free that takes the process down rather than throwing. + after `REAP_IDLE_MS` unused and rebuilt on next use. The window is RAM-SCALED (120 s low-RAM, + 600 s roomy): a reload is ~1 s of dead mic on the next tap, and on a roomy phone that latency + regression buys nothing - pre-reaper those devices held the model all session and were fine. + `WhisperRecognizer.release()` declines while a listen is in flight - freeing the native + recognizer under a running decode is a use-after-free that takes the process down rather than + throwing. + - **A model's file size says NOTHING about its resident cost - measure before adopting.** NeMo + Conformer CTC small is a 46 MB int8 file that ballooned to ~760 MB-1.2 GB PSS through + onnxruntime on the M5 (rejected); Moonshine's small weights still cost ~212 MB across its four + ORT sessions - no lighter resident than Whisper tiny's ~214 MB (both same-protocol launch + deltas, 32-bit M5). Zipformer small (26 MB encoder, one tiny decoder/joiner pair) is the + structural low-memory candidate; its isolated reap-delta number is tracked in PR #86 and MUST + be written here before any low-RAM default is switched to it. "Encoder-only means small" and + "fewer MB means less RAM" were both device-refuted in one afternoon. The clean way to isolate + one model's resident cost: let the IDLE REAP fire (it releases ONLY the recognizer) and diff + `Native Heap` pre/post inside one settled process - launch-to-launch totals swing +-150 MB with + map content and prove nothing at engine granularity. - **Do not reach for a device gate when an IDLE gate will do.** The warm-at-startup behaviour was first made low-RAM-conditional, which protected the instant-first-mic-tap UX on roomier phones but left them holding 267 MB all session. Reaping on idle keeps that UX AND reclaims the memory @@ -1013,17 +1027,48 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): `onTrimMemory` on the main thread, and `loadLock` is held across the ~1 s native model load, so a blocking acquire stalls the UI thread for that whole load just to reclaim memory the idle reaper would reclaim anyway. Skipping is safe; the next trim or the reaper retries. `Remove - model` passes `wait = true` because there the user asked for it. Device-verified that this - window is REAL and not theoretical: hammering `am send-trim-memory RUNNING_CRITICAL` - across startup logged `release skipped, model load in progress` 5 times in one run. + model` passes `wait = true` because there the user asked for it - and `deleteAsrEngine` runs + that whole sequence (wait + native free + the up-to-154 MB recursive delete) on `Dispatchers.IO`, + never the UI thread. Device-verified that this window is REAL and not theoretical: hammering + `am send-trim-memory RUNNING_CRITICAL` across startup logged `release skipped, model load + in progress` 5 times in one run. - `RUNNING_CRITICAL` is the level to use for this - it is `isSevere` AND the OS accepts it on a FOREGROUND process, so it exercises the load window without needing HOME first. - - **A quarantined ASR model makes `warmUp()` a silent no-op.** `WhisperRecognizer.isInstalled()` - is `AsrModel.isInstalled() && !asr_model_bad`, and the corrupt-model quarantine (`asr_model_bad` - in `vela_settings`) is only ever lifted by the installer's `clearQuarantine()`. Side-loading the - model files by hand leaves the flag set, so the model never loads, `scudo:secondary` sits at - ~11 MB instead of ~111 MB, and a memory benchmark silently measures the model-absent case. Check - `secondary` is actually ~111 MB before believing any ASR memory number. + - **The lease protects EVERY path that frees the recognizer, including the engine-switch rebuild.** + `ensureRecognizerLocked`'s key-mismatch branch must not `release()` the superseded model while + `leases > 0` - a listen can still be decoding inside it (mic tap, then Settings "Use" on another + engine). It parks the old recognizer in `retired` instead, drained under `loadLock` once the + lease count returns to zero. Holding two models briefly is the price of not freeing one under a + running decode. + - **Taking a lease across a suspension point must be CANCELLATION-SAFE.** `withContext { acquire() }` + runs the block to completion and then throws `CancellationException` INSTEAD of returning the + value when the caller was cancelled mid-load - so a `finally` keyed off the returned reference + leaks the lease forever (every release path then declines for the rest of the process, and the + 267 MB becomes unreclaimable). `listen()` sets a `leased` flag INSIDE the block via `.also {}` + and keys the `finally` off the flag, not the reference. + - **A trim-triggered WebView reap DECLINES while a fetch is in flight, except at CRITICAL.** All + five fetchers' `reapNow(force)` skip teardown when `pending` is non-empty and the trim is merely + severe - destroying the view kills the injected scraper and turns a live gallery/reviews/ + directions fetch into an empty panel, and the fetch's own `finally` re-arms the reap moments + later anyway. A CRITICAL trim forces the teardown and then DRAINS `pending` (complete-empty), so + the stranded fetch fails fast instead of parking in `deferred.await()` for the full timeout + while holding the serializing mutex. All reap bookkeeping is main-thread-confined via `onMain` + in every fetcher - `reap` scheduling is a read-modify-write and @Volatile does not fix one. + - **The sherpa-onnx AAR is pinned at >= 1.13.4 because of 32-bit ARM.** Its bundled onnxruntime + (1.27.0) fixes an unaligned-read SIGBUS (`BUS_ADRALN`) that crashed every MODEL LOAD on + armeabi-v7a - TTS and ASR both, uncatchably, at startup (issue #95). Device-verified both ways + on an M5 forced to `--abi armeabi-v7a`: 1.13.3/ORT 1.24.3 SIGBUSes in `libonnxruntime.so` on + the loading thread; 1.13.4/ORT 1.27.0 loads and releases clean. The feature phones this fork + exists for have NO system speech engine - the downloaded models are their only voice - so "gate + voice off on 32-bit" is not an acceptable fallback and the runtime must keep working there. + Test any AAR bump on a 32-bit install before shipping it. + - **A quarantined model makes `warmUp()` a silent no-op.** The corrupt-model quarantine keys + (`asr_model_bad_`, `piper_model_bad_` in `vela_settings`) are only ever + lifted by the installer paths' `clearQuarantine()`. Side-loading model files by hand leaves a + latched flag set, so the model never loads, `scudo:secondary` sits at ~11 MB instead of + ~111 MB, and a memory benchmark silently measures the model-absent case. Check `secondary` is + actually model-sized before believing any ASR/TTS memory number - and clear BOTH the bad flag + and the strike counter when resetting a device by hand. ## Performance: what has been measured @@ -1679,7 +1724,13 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): 11 languages; URL = `…/tts-models/vits-piper-.tar.bz2`). `PiperSynth.ensureLoaded` reloads when the selected voice changes; `PiperSynth.reloadVoice()` is the SINGLE switch trigger - it bumps the generation counter (aborting any in-flight utterance) then tears down + rebuilds on the same serial - worker, so `tts` is never freed mid-`generate()`. `MapViewModel.migrateFlatLayoutIfNeeded` (first + worker, so `tts` is never freed mid-`generate()`. The MEMORY-PRESSURE release deliberately does + NOT bump the generation (`release(interrupt = false)`): a trim must reclaim the model, not cut + off the nav prompt being spoken mid-word - the serial worker orders the free after the current + utterance either way. `ensureLoaded` carries the same TWO-STRIKE per-voice crash sentinel as the + ASR loads (`piper_load_strikes_`/`piper_model_bad_`): a voice whose native load + dies (SIGBUS/segfault - uncatchable) crash-looped the app at EVERY launch (issue #95, TCL Flip 2) + because `warmUp()` runs at startup and `catch (Throwable)` never saw the abort. `MapViewModel.migrateFlatLayoutIfNeeded` (first thing in `init`) relocates the old flat single-voice install in place (rename, copy-fallback, verify-gated, re-runnable) - never re-downloads. - **Voice search (speak a query into the search bar), two tiers.** `ui/VoiceSearch` (process-wide @@ -1692,9 +1743,14 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): D-pad-focusable `Dialog`, Done auto-focuses); wiring + the RECORD_AUDIO launcher + the download-offer are in `MapScreen`; the Settings -> Search per-engine picker is in `SettingsScreen`. Needs `RECORD_AUDIO` (manifest; asked at the mic tap). - - **PICKABLE ENGINES via `voice/AsrEngine` (an enum catalog; ported upstream 5d2a6636 / 118e7e8c / - 137beea9).** Three: `WHISPER_TINY` (multilingual, ~58 MB, the `DEFAULT`), `SENSE_VOICE` - (en/zh/ja/ko/yue, ~154 MB), `MOONSHINE` (English-only, ~101 MB). Each is an OPTIONAL download to + - **PICKABLE ENGINES via `voice/AsrEngine` (an enum catalog; first three ported upstream + 5d2a6636 / 118e7e8c / 137beea9).** Four: `WHISPER_TINY` (multilingual, ~58 MB, the `DEFAULT`), + `SENSE_VOICE` (en/zh/ja/ko/yue, ~154 MB), `MOONSHINE` (English-only, ~101 MB), + `ZIPFORMER_SMALL` (English-only, ~28 MB, this fork's low-memory pick - and it emits ALL-CAPS + spoken-form text, so `SpeechText.cleanSearchTranscript` lowercases and runs + `spokenNumbersToDigits`, the unit-tested inverse text normalization that turns "ONE TWENTY + THREE MAIN STREET" into "123 main street"; Whisper writes digits itself and passes through + unchanged). Each is an OPTIONAL download to `filesDir/asr//`; `active()` is the picked engine (pref `asr_engine`), `forRecognition(lang)` falls back to Whisper when the pick can't do the app language. **Whisper stays the default and the ONLY thing onboarding / the map mic offer install** (`downloadAsrModel()` = diff --git a/app/build.gradle.kts b/app/build.gradle.kts index 2181c57d..7b256c64 100644 --- a/app/build.gradle.kts +++ b/app/build.gradle.kts @@ -276,11 +276,13 @@ dependencies { implementation(project(":core")) implementation(project(":yapchik")) // vendored softkey engine (LGPL-3.0) - keypad/D-pad softkeys - // sherpa-onnx: in-process neural TTS runtime (runs the downloaded Kokoro model). Vendored AAR - // (no official Maven artifact; the JitPack coordinate doesn't resolve). Lives in :app because a - // library module can't consume a local .aar - KokoroSynth sits in :app and bridges into :core's - // VoiceGuide via an interface. Native .so are arm64-only in the package (see packaging{}). - implementation(files("libs/sherpa-onnx-1.13.3.aar")) + // sherpa-onnx: in-process neural TTS + ASR runtime. Vendored AAR (no official Maven artifact; + // the JitPack coordinate doesn't resolve). Lives in :app because a library module can't consume + // a local .aar. 1.13.4 is a LOAD-BEARING upgrade, not routine: its bundled onnxruntime (1.27.0, + // up from 1.24.3) fixes the armv7 unaligned-read SIGBUS that crashed every model LOAD on 32-bit + // ARM phones (issue #95; device-verified both broken-before and fixed-after on an M5 forced to + // `--abi armeabi-v7a`). Do not downgrade past it while the fork ships v7a. + implementation(files("libs/sherpa-onnx-1.13.4.aar")) // Extracts the Kokoro model's .tar.bz2 at download time (Android has no built-in bzip2/tar). implementation("org.apache.commons:commons-compress:1.27.1") diff --git a/app/src/main/java/app/vela/data/FlockCameras.kt b/app/src/main/java/app/vela/data/FlockCameras.kt index 8ed0d59f..56f8b021 100644 --- a/app/src/main/java/app/vela/data/FlockCameras.kt +++ b/app/src/main/java/app/vela/data/FlockCameras.kt @@ -34,48 +34,60 @@ object FlockCameras { private const val BUNDLED_VER = "flock_cameras_version.txt" private const val CELL = 0.1 // grid cell size in degrees (~11 km) for the bucket index - @Volatile private var loaded = false - private var lat = DoubleArray(0) - private var lng = DoubleArray(0) - private var op = arrayOf() /** - * The 0.1 deg bucket index, as a flat CSR (compressed sparse row) triple rather than a - * `HashMap>`. + * One GENERATION of the dataset: coordinates, operators, and the 0.1 deg bucket index as a flat + * CSR (compressed sparse row) triple rather than a `HashMap>`. * * [cellKeys] is sorted and unique; the rows in cell `k` are * `cellRows[cellStart[k] until cellStart[k + 1]]`. Lookup is a binary search on [cellKeys]. * - * Why: the map version cost roughly 166,000 objects to build - one boxed `Integer` per camera + * Why CSR: the map version cost roughly 166,000 objects to build - one boxed `Integer` per camera * (124,406 of them), plus an `ArrayList` + its `Object[]` + a boxed `Long` key + a `HashMap.Node` * per occupied cell (13,965 of those) - and every one of them was garbage the moment the parse * finished. This is five arrays. Startup on this app is GC-bound, not compilation-bound (forcing * full AOT made cold start WORSE, 828 ms against 775 ms), so allocation count at startup is the * thing worth cutting. Lookups are unchanged in behaviour and no slower in practice: the viewport * scan touches on the order of 100 cells, and a binary search over ~14,000 keys is ~14 compares. + * + * Why one class and not six fields on the object: the six arrays are one INVARIANT - a lookup's + * key array and offset array must come from the same parse. `refresh()`'s hot-swap runs on + * Dispatchers.IO while a viewport scan may be walking the index, and separately-published fields + * can tear (new, larger `cellKeys` against old `cellStart` = `cellStart[k + 1]` past the end - + * an ArrayIndexOutOfBoundsException out of `inBox`, which nothing catches). A single @Volatile + * reference snapshotted once per query cannot: a reader sees the whole old generation or the + * whole new one. */ - private var cellKeys = LongArray(0) - private var cellStart = IntArray(1) - private var cellRows = IntArray(0) + private class Data( + val lat: DoubleArray, + val lng: DoubleArray, + val op: Array, + val cellKeys: LongArray, + val cellStart: IntArray, + val cellRows: IntArray, + ) { + /** + * Run [body] for every camera row in the cell at ([row], [col]). No-op when the cell is + * empty. Replaces `grid[key(r, c)]?.let { ... }`, and allocates nothing - notably no + * iterator, which the old `for (i in bucket)` over a `MutableList` created (and + * unboxed on every step). + */ + inline fun forEachInCell(row: Long, col: Long, body: (Int) -> Unit) { + val k = cellKeys.binarySearch(key(row, col)) + if (k < 0) return + var p = cellStart[k] + val end = cellStart[k + 1] + while (p < end) { body(cellRows[p]); p++ } + } + } - val isLoaded: Boolean get() = loaded - val size: Int get() = lat.size + @Volatile private var data: Data? = null + + val isLoaded: Boolean get() = data != null + val size: Int get() = data?.lat?.size ?: 0 private fun key(row: Long, col: Long): Long = (row shl 32) xor (col and 0xffffffffL) private fun rowOf(v: Double): Long = Math.floor(v / CELL).toLong() - /** - * Run [body] for every camera row in the cell at ([row], [col]). No-op when the cell is empty. - * Replaces `grid[key(r, c)]?.let { ... }`, and allocates nothing - notably no iterator, which - * the old `for (i in bucket)` over a `MutableList` created (and unboxed on every step). - */ - private inline fun forEachInCell(row: Long, col: Long, body: (Int) -> Unit) { - val k = cellKeys.binarySearch(key(row, col)) - if (k < 0) return - var p = cellStart[k] - val end = cellStart[k + 1] - while (p < end) { body(cellRows[p]); p++ } - } - private fun dir(context: Context) = File(context.filesDir, "flock").apply { mkdirs() } private fun downloadedBin(context: Context) = File(dir(context), "cameras.bin") private fun downloadedVer(context: Context) = File(dir(context), "version.txt") @@ -90,9 +102,9 @@ object FlockCameras { /** Parse the newest available file once, off the main thread. Safe to call repeatedly (a loaded call no-ops). */ suspend fun ensureLoaded(context: Context) { - if (loaded) return + if (data != null) return withContext(Dispatchers.IO) { - if (loaded) return@withContext + if (data != null) return@withContext val dl = downloadedBin(context) val stream = if (dl.exists()) runCatching { dl.inputStream() }.getOrNull() else runCatching { context.assets.open(BUNDLED) }.getOrNull() @@ -147,14 +159,20 @@ object FlockCameras { val rows = IntArray(n) for (i in 0 until n) { val k = ck.binarySearch(keys[i]); rows[fill[k]] = i; fill[k]++ } - // Publish (a bad/partial parse threw before here, so we never swap in a half-built set). - lat = las.copyOf(n); lng = los.copyOf(n) + // Publish the whole generation in ONE volatile write (a bad/partial parse threw before + // here, so we never swap in a half-built set - and a concurrent reader holding the old + // generation's snapshot keeps using it, consistently, until its query ends). @Suppress("UNCHECKED_CAST") - op = (ops.copyOf(n) as Array) - cellKeys = ck; cellStart = start; cellRows = rows - loaded = true + data = Data(las.copyOf(n), los.copyOf(n), ops.copyOf(n) as Array, ck, start, rows) } + /** Test seam: build a generation from an in-memory stream (unit tests have no Context/assets). */ + @androidx.annotation.VisibleForTesting + internal fun loadFromForTest(raw: InputStream) = loadFrom(raw) + + @androidx.annotation.VisibleForTesting + internal fun resetForTest() { data = null } + private val downloadHttp: OkHttpClient by lazy { OkHttpClient.Builder().callTimeout(0, TimeUnit.SECONDS).readTimeout(60, TimeUnit.SECONDS).build() } @@ -191,7 +209,7 @@ object FlockCameras { /** Cameras inside the bbox, for DRAWING. Empty if not loaded yet (caller falls back to Overpass). */ fun inBox(south: Double, west: Double, north: Double, east: Double): List { - if (!loaded) return emptyList() + val d = data ?: return emptyList() // ONE snapshot per query - see [Data] val out = ArrayList() val r0 = rowOf(south); val r1 = rowOf(north) val c0 = rowOf(west); val c1 = rowOf(east) @@ -199,8 +217,10 @@ object FlockCameras { while (r <= r1) { var c = c0 while (c <= c1) { - forEachInCell(r, c) { i -> - if (lat[i] in south..north && lng[i] in west..east) out.add(AlprCamera(LatLng(lat[i], lng[i]), op[i])) + d.forEachInCell(r, c) { i -> + if (d.lat[i] in south..north && d.lng[i] in west..east) { + out.add(AlprCamera(LatLng(d.lat[i], d.lng[i]), d.op[i])) + } } c++ } @@ -211,7 +231,8 @@ object FlockCameras { /** Cameras within [meters] of any SEGMENT of [polyline], for the route count. Empty if not loaded. */ fun along(polyline: List, meters: Double = 120.0): List { - if (!loaded || polyline.size < 2) return emptyList() + val d = data ?: return emptyList() // ONE snapshot per query - see [Data] + if (polyline.size < 2) return emptyList() val pad = 0.01 val r0 = rowOf(polyline.minOf { it.lat } - pad); val r1 = rowOf(polyline.maxOf { it.lat } + pad) val c0 = rowOf(polyline.minOf { it.lng } - pad); val c1 = rowOf(polyline.maxOf { it.lng } + pad) @@ -220,9 +241,9 @@ object FlockCameras { while (r <= r1) { var c = c0 while (c <= c1) { - forEachInCell(r, c) { i -> - val p = LatLng(lat[i], lng[i]) - if (nearPolyline(p, polyline, meters)) out.add(AlprCamera(p, op[i])) + d.forEachInCell(r, c) { i -> + val p = LatLng(d.lat[i], d.lng[i]) + if (nearPolyline(p, polyline, meters)) out.add(AlprCamera(p, d.op[i])) } c++ } diff --git a/app/src/main/java/app/vela/ui/map/MapViewModel.kt b/app/src/main/java/app/vela/ui/map/MapViewModel.kt index 0ceb517e..affecdf5 100644 --- a/app/src/main/java/app/vela/ui/map/MapViewModel.kt +++ b/app/src/main/java/app/vela/ui/map/MapViewModel.kt @@ -3472,6 +3472,7 @@ class MapViewModel @Inject constructor( ) } if (ok && VelaPiper.isVoiceReady(appContext, id)) { + piperSynth.clearQuarantine(id) // a fresh download replaces whatever was quarantined if (firstEver) selectVoice(id) else flashStatus(appContext.getString(R.string.mapvm_voice_downloaded, v.displayName)) } else { showStatus(appContext.getString(R.string.mapvm_voice_download_failed, v.displayName)) @@ -4474,13 +4475,20 @@ class MapViewModel @Inject constructor( /** Remove one engine's model (Settings "Remove"); the active pick degrades to another installed * engine automatically (AsrEngine.active). */ fun deleteAsrEngine(engine: app.vela.voice.AsrEngine) { - // Free the loaded model BEFORE removing its files. Deleting the directory alone left the - // native recognizer resident for the rest of the process (~267 MB measured, issue #83), so - // "Remove" reclaimed disk but no memory at all. Released unconditionally: the loaded engine - // may not be [engine], but the worst case is a ~1 s reload on the next listen. - whisperRecognizer.release(wait = true) // deliberate user action: worth waiting out a load - engine.dir(appContext).deleteRecursively() - refreshAsr() + // Drop it from the UI immediately (optimistic, same idiom as deleteVoice); the real work is + // OFF the main thread: release(wait = true) can park behind a multi-second in-flight native + // load on loadLock, then frees ~267 MB of native memory, and the recursive delete unlinks up + // to 154 MB of files - all three are ANR material on a slow keypad phone's UI thread. + _state.update { it.copy(asrInstalledIds = it.asrInstalledIds - engine.id) } + viewModelScope.launch(kotlinx.coroutines.Dispatchers.IO) { + // Free the loaded model BEFORE removing its files. Deleting the directory alone left the + // native recognizer resident for the rest of the process (~267 MB measured, issue #83), + // so "Remove" reclaimed disk but no memory at all. Released unconditionally: the loaded + // engine may not be [engine], but the worst case is a ~1 s reload on the next listen. + whisperRecognizer.release(wait = true) // deliberate user action: worth waiting out a load + engine.dir(appContext).deleteRecursively() + refreshAsr() + } } /** Onboarding offers BOTH on-device speech models on one screen, so a user can pick both at once. diff --git a/app/src/main/java/app/vela/ui/settings/sections/SearchSettings.kt b/app/src/main/java/app/vela/ui/settings/sections/SearchSettings.kt index aa4727fd..e4d5d7d8 100644 --- a/app/src/main/java/app/vela/ui/settings/sections/SearchSettings.kt +++ b/app/src/main/java/app/vela/ui/settings/sections/SearchSettings.kt @@ -80,6 +80,7 @@ internal fun SearchSettingsScreen(vm: MapViewModel, onBack: () -> Unit) { app.vela.voice.AsrEngine.WHISPER_TINY -> R.string.settings_asr_langs_whisper app.vela.voice.AsrEngine.SENSE_VOICE -> R.string.settings_asr_langs_sensevoice app.vela.voice.AsrEngine.MOONSHINE -> R.string.settings_asr_langs_moonshine + app.vela.voice.AsrEngine.ZIPFORMER_SMALL -> R.string.settings_asr_langs_zipformer }, ) val meta = stringResource(R.string.settings_asr_engine_meta, langs, engine.sizeMb) diff --git a/app/src/main/java/app/vela/voice/AsrEngine.kt b/app/src/main/java/app/vela/voice/AsrEngine.kt index 77f301ea..e7afa894 100644 --- a/app/src/main/java/app/vela/voice/AsrEngine.kt +++ b/app/src/main/java/app/vela/voice/AsrEngine.kt @@ -17,15 +17,26 @@ private const val VAD_FILE = "silero_vad.onnx" * no account or third-party voice app is needed. Each engine is an OPTIONAL one-time download hosted * on this repo's `asr-models` GitHub release, extracted to `filesDir/asr//`. * - * Three engines, because they trade off differently and the user picks (ported from upstream - * PimpinPumpkin/Vela 5d2a6636 + 118e7e8c sizes + 137beea9 language fallback): + * Four engines, because they trade off differently and the user picks (the first three ported from + * upstream PimpinPumpkin/Vela 5d2a6636 + 118e7e8c sizes + 137beea9 language fallback): * - [WHISPER_TINY] - the multilingual default. 99-language Whisper tiny (int8); covers every - * language Vela's UI supports (incl. Hebrew, Russian, Spanish). The safe all-rounder, and the - * smallest - so it stays the default and the ONLY thing the one-tap onboarding/map offer installs, - * which matters on the RAM/storage-constrained feature phones this fork targets. + * language Vela's UI supports (incl. Hebrew, Russian, Spanish). The safe all-rounder and the + * smallest download - so it stays the default and the ONLY thing the one-tap onboarding/map + * offer installs. NOT the smallest loaded: ~214 MB PSS resident (measured, 32-bit M5). * - [SENSE_VOICE] - FunAudioLLM SenseVoice. More accurate + faster than Whisper tiny, but only for * English, Chinese, Cantonese, Japanese, Korean. Bigger (opt-in). - * - [MOONSHINE] - Useful Sensors Moonshine tiny. Lowest latency, ENGLISH ONLY. Bigger (opt-in). + * - [MOONSHINE] - Useful Sensors Moonshine tiny. Lowest latency, ENGLISH ONLY. Bigger (opt-in), + * and despite the small weights its four ORT sessions cost ~212 MB PSS loaded - no lighter + * resident than Whisper (measured, 32-bit M5). + * - [ZIPFORMER_SMALL] - k2 Zipformer small transducer, ENGLISH ONLY. The low-memory candidate for + * the RAM-constrained feature phones this fork exists for (this fork's addition): 26 MB int8 + * encoder + kilobyte-scale decoder/joiner, no 30 s fixed context and no KV cache. Transducer + * over CTC deliberately - sherpa-onnx hotword/contextual biasing (street/place names) works on + * transducers, and the zoo has no offline English CTC small anyway. NeMo Conformer CTC small + * was tried first for this slot and REJECTED on measurement: its int8 graph ballooned to + * ~760 MB-1.2 GB PSS through onnxruntime on the M5, worse than Whisper - "encoder-only means + * small" did not survive a device number, so do not trust this entry's resident cost until the + * isolated reap-delta measurement in PR #86 is written back here. * * Whisper stays the default so no language silently regresses; the other two are opt-in via the * voice-search engine picker in Settings. This holds only metadata + a cheap install check + the @@ -71,6 +82,14 @@ enum class AsrEngine( "cached_decode.int8.onnx", "tokens.txt", VAD_FILE, ), ), + ZIPFORMER_SMALL( + id = "zipformer-small-en", + displayName = "Zipformer small", + modelType = "transducer", + sizeMb = 28, + url = "$ASR_BASE/vela-asr-zipformer-small-en.tar.bz2", + files = listOf("encoder.int8.onnx", "decoder.int8.onnx", "joiner.int8.onnx", "tokens.txt", VAD_FILE), + ), ; /** `filesDir/asr//` - the extracted archive's single top-level folder. */ @@ -90,6 +109,7 @@ enum class AsrEngine( WHISPER_TINY -> true SENSE_VOICE -> lang in SENSE_VOICE_LANGS MOONSHINE -> lang == "en" + ZIPFORMER_SMALL -> lang == "en" } companion object { diff --git a/app/src/main/java/app/vela/voice/PiperSynth.kt b/app/src/main/java/app/vela/voice/PiperSynth.kt index 99f30f01..fa1dcf36 100644 --- a/app/src/main/java/app/vela/voice/PiperSynth.kt +++ b/app/src/main/java/app/vela/voice/PiperSynth.kt @@ -44,10 +44,14 @@ class PiperSynth @Inject constructor( // synth costs a reload on the next prompt, and a prompt arriving late during navigation is // a missed turn. TRIM_MEMORY_COMPLETE only reaches background processes, and // RUNNING_CRITICAL means the device is about to start killing things regardless. - // release() posts to the piper-tts worker, so it is already serialized against an + // The teardown posts to the piper-tts worker, so it is already serialized against an // in-flight synthesis and cannot free the model out from under one (issue #83). + // interrupt = false: a trim must reclaim memory, not SILENCE the prompt being spoken - + // release()'s default generation bump aborts an in-flight utterance within ~200 ms, which + // during navigation is a missed turn (the very regression scoping to CRITICAL was meant to + // avoid). Un-bumped, the serial worker frees the model right AFTER the current utterance. app.vela.ui.MemoryPressure.register { level -> - if (app.vela.ui.MemoryPressure.isCritical(level)) release() + if (app.vela.ui.MemoryPressure.isCritical(level)) release(interrupt = false) } } @@ -92,6 +96,16 @@ class PiperSynth @Inject constructor( worker.execute { ensureLoaded() } } + private fun prefs() = context.getSharedPreferences("vela_settings", Context.MODE_PRIVATE) + + /** Lift a voice's crash quarantine after a fresh download - the bad files are gone, so the next + * load may try again. Called by the installer path, never automatically. */ + fun clearQuarantine(voiceId: String) { + prefs().edit() + .putBoolean(KEY_MODEL_BAD + voiceId, false) + .putInt(KEY_LOAD_STRIKES + voiceId, 0).apply() + } + private fun ensureLoaded(): OfflineTts? { val r = VelaPiper.resolved(context) ?: return null // nothing usable installed val cur = tts @@ -101,10 +115,32 @@ class PiperSynth @Inject constructor( // use-after-free). loadFailed resets so a previously-bad voice doesn't block a new one. runCatching { cur?.release() } tts = null; loadedVoiceId = null; numSpeakers = 0; loadFailed = false + // CRASH SENTINEL around the native load, PER VOICE - the same two-strike idiom as + // WhisperRecognizer's ASR loads, closing the same hole for TTS (issue #95: a voice whose + // load dies natively - SIGBUS/segfault, which no `catch (Throwable)` can see - crash-looped + // the app at EVERY launch, because warmUp() runs at startup and nothing remembered the + // previous attempt never returned). Bump a strike before the load, zero it once it returns; + // two stranded loads in a row quarantine THAT voice only and delete its dir, so the app + // boots (system TTS takes over) and a fresh download starts clean. Two, not one: a process + // killed mid-load (swipe-away, memory reclaim) strands a strike exactly like a crash, and + // must not delete a healthy 80 MB voice. + val prefs = prefs() + val strikesKey = KEY_LOAD_STRIKES + r.voiceId + val badKey = KEY_MODEL_BAD + r.voiceId + val strikes = prefs.getInt(strikesKey, 0) + if (strikes >= 2) { + Timber.tag(TAG).e("two voice loads never returned (native crash) - quarantining ${r.voiceId}") + prefs.edit().putInt(strikesKey, 0).putBoolean(badKey, true).apply() + runCatching { VelaPiper.modelDirFor(context, r.voiceId).deleteRecursively() } + loadFailed = true + return null + } + if (prefs.getBoolean(badKey, false)) { loadFailed = true; return null } // Two attempts: a voice loaded the instant its download/extract finishes can lose the race with // the filesystem flush on some devices - the first OfflineTts load throws, and (without a retry) // loadFailed sticks so the voice stays SILENT until an app restart. A brief retry heals it. repeat(2) { attempt -> + prefs.edit().putInt(strikesKey, prefs.getInt(strikesKey, 0) + 1).apply() try { // Lower the VITS noise scales below the library defaults (noiseScale 0.667, noiseScaleW 0.8). // Those defaults make synthesis STOCHASTIC - the same phrase varies run to run, which is why @@ -123,11 +159,15 @@ class PiperSynth @Inject constructor( val engine = OfflineTts(assetManager = null, config = cfg) numSpeakers = engine.numSpeakers() runCatching { engine.generate(text = " ", sid = 0, speed = SPEED) } + // The load (and warm synth) RETURNED - the process survived it, so zero the strikes. + // Only a native abort mid-load leaves one standing. + prefs.edit().putInt(strikesKey, 0).apply() tts = engine loadedVoiceId = r.voiceId Timber.tag(TAG).i("loaded ${r.voiceId}: sampleRate=${engine.sampleRate()} speakers=$numSpeakers") return engine } catch (t: Throwable) { + prefs.edit().putInt(strikesKey, 0).apply() // a CATCHABLE failure is not a native crash Timber.tag(TAG).e(t, "model load failed (attempt ${attempt + 1}): ${t.message}") if (attempt == 0) runCatching { Thread.sleep(200) } // let a just-written model settle, then retry } @@ -296,8 +336,13 @@ class PiperSynth @Inject constructor( worker.execute { runCatching { track?.pause(); track?.flush() } } } - override fun release() { - generation++ + override fun release() = release(interrupt = true) + + /** Free the engine + track. [interrupt] aborts any in-flight utterance first (the right thing + * when the voice is being switched off or replaced); the memory-pressure path passes false so + * the current prompt finishes - the serial worker orders the free after it either way. */ + fun release(interrupt: Boolean) { + if (interrupt) generation++ worker.execute { runCatching { track?.release() }; track = null runCatching { tts?.release() }; tts = null @@ -307,6 +352,10 @@ class PiperSynth @Inject constructor( private companion object { const val TAG = "PiperSynth" const val SPEED = 1.0f + // Crash-sentinel keys, PER VOICE (suffixed with the voice id) - same idiom and reasoning as + // WhisperRecognizer's KEY_LOAD_STRIKES/KEY_MODEL_BAD, see the sentinel in [ensureLoaded]. + const val KEY_LOAD_STRIKES = "piper_load_strikes_" + const val KEY_MODEL_BAD = "piper_model_bad_" // Silence spliced between sentences (seconds) - a natural period beat for nav prompts. const val PAUSE_SEC = 0.32f // Shorter beat spliced at commas/semicolons so clauses don't run together ("In a quarter mile, …"). diff --git a/app/src/main/java/app/vela/voice/WhisperRecognizer.kt b/app/src/main/java/app/vela/voice/WhisperRecognizer.kt index b9cb65c9..b58f2f2a 100644 --- a/app/src/main/java/app/vela/voice/WhisperRecognizer.kt +++ b/app/src/main/java/app/vela/voice/WhisperRecognizer.kt @@ -14,6 +14,7 @@ import timber.log.Timber import com.k2fsa.sherpa.onnx.FeatureConfig import com.k2fsa.sherpa.onnx.OfflineModelConfig import com.k2fsa.sherpa.onnx.OfflineMoonshineModelConfig +import com.k2fsa.sherpa.onnx.OfflineTransducerModelConfig import com.k2fsa.sherpa.onnx.OfflineSenseVoiceModelConfig import com.k2fsa.sherpa.onnx.OfflineRecognizer import com.k2fsa.sherpa.onnx.OfflineRecognizerConfig @@ -68,6 +69,14 @@ class WhisperRecognizer @Inject constructor( */ private val leases = java.util.concurrent.atomic.AtomicInteger(0) + /** Recognizers superseded by an engine/language switch WHILE a lease was outstanding. The + * switch path must not free the old engine then - the lease-holding decode is still inside + * it, and `OfflineRecognizer.release()` frees C++ memory (the same use-after-free [leases] + * exists to stop). Parked here instead, freed by [drainRetiredLocked] once every lease is + * back. Guarded by [loadLock]. Briefly costs two resident models; a switch mid-listen is + * rare enough that correctness wins. */ + private val retired = ArrayList() + /** Idle-reap timer, same idea as the web fetchers' `REAP_IDLE_MS` (issue #182). One daemon * thread, shared, created lazily so a device that never loads the model never starts it. */ private val reaper by lazy { @@ -127,6 +136,7 @@ class WhisperRecognizer @Inject constructor( Timber.tag(TAG).i("release skipped, listen in flight") return } + drainRetiredLocked() val r = recognizer ?: return recognizer = null loadedKey = null @@ -158,13 +168,33 @@ class WhisperRecognizer @Inject constructor( } /** - * Give back a lease taken by [acquireRecognizer]. Deliberately does NOT take [loadLock]: the + * Give back a lease taken by [acquireRecognizer]. Deliberately does NOT block on [loadLock]: the * decrement happens only once the decode is finished with the pointer, so the worst a racing * [release] can do is read the pre-decrement value and conservatively decline. Taking the lock - * here would instead park the end of every utterance behind an unrelated model load. + * here would instead park the end of every utterance behind an unrelated model load - so the + * retired-model drain runs only when the lock is free, and every other lock-holder (acquire, + * release, the next load) drains as well, so a skipped drain is picked up at the next one. */ private fun releaseLease() { - leases.decrementAndGet() + if (leases.decrementAndGet() == 0 && loadLock.tryLock()) { + try { + drainRetiredLocked() + } finally { + loadLock.unlock() + } + } + } + + /** Free every recognizer parked by an engine switch, once no lease can still be inside one. + * **Caller must hold [loadLock].** */ + private fun drainRetiredLocked() { + if (leases.get() > 0 || retired.isEmpty()) return + for (r in retired) { + runCatching { r.release() } + .onFailure { Timber.tag(TAG).w(it, "retired recognizer release failed") } + } + retired.clear() + Timber.tag(TAG).i("retired recognizer(s) released") } private val audioManager by lazy { context.getSystemService(Context.AUDIO_SERVICE) as? AudioManager } @@ -209,10 +239,13 @@ class WhisperRecognizer @Inject constructor( const val SAMPLE_RATE = 16000 const val VAD_WINDOW = 512 // Silero v4/v5 window at 16 kHz const val MAX_SECONDS = 15 // hard cap on one utterance - // Drop the loaded model after this quiet period (issue #83). Matches the web - // fetchers' REAP_IDLE_MS: long enough that a dictation session never reaps between - // utterances, short enough that a session-long 267 MB hold cannot happen. - const val REAP_IDLE_MS = 120_000L + // Drop the loaded model after this quiet period (issue #83): long enough that a dictation + // session never reaps between utterances, short enough that a session-long 267 MB hold + // cannot happen. RAM-SCALED, not flat: a reload costs ~1 s of dead mic on the next tap, + // and on a roomy phone that latency regression buys nothing the phone needed - pre-#83 + // those devices held the model all session and were fine. 2 min where the 267 MB actually + // hurts, 10 min where it is merely tidy. + val REAP_IDLE_MS: Long get() = if (app.vela.ui.MemoryPressure.lowRam) 120_000L else 600_000L } private fun prefs() = context.getSharedPreferences("vela_settings", Context.MODE_PRIVATE) @@ -252,6 +285,7 @@ class WhisperRecognizer @Inject constructor( AsrEngine.WHISPER_TINY -> l.takeIf { it in app.vela.ui.AppLocale.SUPPORTED } ?: "" AsrEngine.SENSE_VOICE -> l.takeIf { it in AsrEngine.SENSE_VOICE_LANGS } ?: "auto" AsrEngine.MOONSHINE -> "" + AsrEngine.ZIPFORMER_SMALL -> "" // English-only transducer, takes no language } } @@ -296,10 +330,17 @@ class WhisperRecognizer @Inject constructor( * utterance, and it is what makes the lease in [acquireRecognizer] atomic. */ private fun ensureRecognizerLocked(): OfflineRecognizer? { + drainRetiredLocked() val engine = engineForNow() val lang = pinnedLang(engine) val key = "${engine.id}|$lang" - recognizer?.let { if (loadedKey == key) return it else runCatching { it.release() } } + recognizer?.let { + if (loadedKey == key) return it + // Engine/language switch. Free the superseded model only if no listen can still be + // decoding inside it; otherwise park it in [retired] - releasing native memory under + // an outstanding lease is the use-after-free the lease exists to prevent. + if (leases.get() > 0) retired.add(it) else runCatching { it.release() } + } recognizer = null if (!engine.isInstalled(context)) return null @@ -367,6 +408,16 @@ class WhisperRecognizer @Inject constructor( numThreads = 2, modelType = engine.modelType, ) + AsrEngine.ZIPFORMER_SMALL -> OfflineModelConfig( + transducer = OfflineTransducerModelConfig( + encoder = p("encoder.int8.onnx"), + decoder = p("decoder.int8.onnx"), + joiner = p("joiner.int8.onnx"), + ), + tokens = p("tokens.txt"), + numThreads = 2, + modelType = engine.modelType, + ) } val r = runCatching { OfflineRecognizer( @@ -416,17 +467,23 @@ class WhisperRecognizer @Inject constructor( // listen() from a UI coroutine. The lease is taken here rather than inside listenInner so // that the many early returns in there cannot leak it - the finally below always gives it // back, and it covers recording as well as decoding. + // + // `leased` is set INSIDE the withContext block, not inferred from `rec`: a coroutine + // cancelled during the ~1 s load makes withContext run the block to completion (taking the + // lease) and then throw CancellationException INSTEAD of returning the value - `rec` would + // never be assigned, and a rec-based finally would leak the lease forever, permanently + // disabling every release path (reaper, trims, Remove model). reapTask?.cancel(false) // never reap mid-utterance - val rec = withContext(Dispatchers.Default) { acquireRecognizer() } - ?: run { - Timber.tag(TAG).e("listen failed: MODEL (model absent or native load failed)") - armIdleReap() - return VoiceResult.Failed(VoiceResult.Reason.MODEL, "model absent or native load failed") - } + var leased = false try { + val rec = withContext(Dispatchers.Default) { acquireRecognizer()?.also { leased = true } } + ?: run { + Timber.tag(TAG).e("listen failed: MODEL (model absent or native load failed)") + return VoiceResult.Failed(VoiceResult.Reason.MODEL, "model absent or native load failed") + } return listenInner(rec, onLevel, onListening, cancelled) } finally { - releaseLease() + if (leased) releaseLease() armIdleReap() // restart the quiet period from the END of this utterance } } diff --git a/app/src/main/java/app/vela/web/WebDirectionsFetcher.kt b/app/src/main/java/app/vela/web/WebDirectionsFetcher.kt index fbf2b053..e3d7883c 100644 --- a/app/src/main/java/app/vela/web/WebDirectionsFetcher.kt +++ b/app/src/main/java/app/vela/web/WebDirectionsFetcher.kt @@ -59,32 +59,54 @@ class WebDirectionsFetcher @Inject constructor( @Volatile private var webView: WebView? = null private var reap: Runnable? = null + /** Run [block] on the main thread, inline when already there. All reap bookkeeping goes through + * here: [reap] is touched by the trim listener (main) and by the fetch path (the caller's + * dispatcher), and scheduling is a read-modify-write - main-confinement removes the race by + * construction (WebPhotoFetcher's fix, issue #83 follow-up). */ + private fun onMain(block: () -> Unit) { + if (Looper.myLooper() == Looper.getMainLooper()) block() else main.post(block) + } + /** Free the WebView after a quiet period (issue #182): a warm fetcher pins a full * maps.google.com page for the rest of the session, and several warm fetchers at once is * real memory. The next fetch after a reap just re-creates it. */ - private fun scheduleReap() { + private fun scheduleReap() = onMain { reap?.let(main::removeCallbacks) - val r = Runnable { reapNow() } + val r = Runnable { reap = null; reapNow() } reap = r main.postDelayed(r, REAP_IDLE_MS) } - /** Destroy the WebView immediately. Must run on the main thread (WebView requirement). */ - private fun reapNow() { + /** Destroy the WebView. Main thread only (WebView requirement). In-flight aware, both ways: + * with a fetch in flight ([pending] non-empty) a non-forced reap DECLINES - destroying the + * view kills the injected scraper and turns a live fetch into an empty result, and the fetch's + * own finally re-arms the reap moments later anyway. A FORCED reap (critical trim - the OS is + * about to start killing) destroys regardless, then drains [pending] so the stranded fetch + * fails fast as empty instead of parking in `deferred.await()` for the full [TOTAL_TIMEOUT_MS] + * while holding [mutex]. */ + private fun reapNow(force: Boolean = false) { + if (!force && pending.isNotEmpty()) { + android.util.Log.i("WebDirectionsFetcher", "reap declined, fetch in flight") + return + } webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } webView = null + pending.keys.toList().forEach { id -> pending.remove(id)?.complete("") } } init { // Under real memory pressure the 120 s idle timer is far too slow - the OS is asking for // memory NOW and a Chromium renderer is one of the largest things we hold (issue #83). - // Reap on the main thread, since WebView.destroy() requires it. + // Reap on the main thread, since WebView.destroy() requires it. Only a CRITICAL trim tears + // down mid-fetch; a merely-severe one declines while a fetch is in flight (see reapNow). app.vela.ui.MemoryPressure.register { level -> - if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { cancelReap(); reapNow() } + if (app.vela.ui.MemoryPressure.isSevere(level)) { + main.post { cancelReap(); reapNow(force = app.vela.ui.MemoryPressure.isCritical(level)) } + } } } - private fun cancelReap() { + private fun cancelReap() = onMain { reap?.let(main::removeCallbacks) reap = null } diff --git a/app/src/main/java/app/vela/web/WebPhotoFetcher.kt b/app/src/main/java/app/vela/web/WebPhotoFetcher.kt index 13e184aa..24a0856b 100644 --- a/app/src/main/java/app/vela/web/WebPhotoFetcher.kt +++ b/app/src/main/java/app/vela/web/WebPhotoFetcher.kt @@ -59,8 +59,12 @@ class WebPhotoFetcher @Inject constructor( init { // A trim is the OS asking for memory NOW, far sooner than any idle timer (issue #83). + // Only a CRITICAL trim tears down mid-scrape; a merely-severe one declines while a fetch + // is in flight (see reapNow) so a moderate-pressure moment doesn't cost a whole gallery. app.vela.ui.MemoryPressure.register { level -> - if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { cancelReap(); reapNow() } + if (app.vela.ui.MemoryPressure.isSevere(level)) { + main.post { cancelReap(); reapNow(force = app.vela.ui.MemoryPressure.isCritical(level)) } + } } } @@ -105,12 +109,19 @@ class WebPhotoFetcher @Inject constructor( /** Destroy the WebView immediately. Main thread only (WebView requirement). The next * [warm]/fetch rebuilds it via `ensureWebView`, exactly as after a renderer death. * - * Drains [pending] for the same reason [rendererGone] does: the injected scraper dies with the - * view, so nothing will ever complete those deferreds. Without this a reap landing mid-fetch - * leaves the fetch parked in `deferred.await()` for the full [TOTAL_TIMEOUT_MS] while it HOLDS - * [mutex], stalling every queued gallery behind it. An empty result is the documented - * best-effort failure mode; a 40 s hang is not. */ - private fun reapNow() { + * A NON-FORCED reap declines while a scrape is in flight ([pending] non-empty): destroying the + * view mid-scrape turns a live gallery into 0 photos, and the fetch's finally re-arms the reap + * moments later anyway. A forced reap (critical trim) proceeds and drains [pending] for the + * same reason [rendererGone] does: the injected scraper dies with the view, so nothing will + * ever complete those deferreds. Without the drain a teardown mid-fetch leaves the fetch + * parked in `deferred.await()` for the full [TOTAL_TIMEOUT_MS] while it HOLDS [mutex], + * stalling every queued gallery behind it. An empty result is the documented best-effort + * failure mode; a 40 s hang is not. */ + private fun reapNow(force: Boolean = false) { + if (!force && pending.isNotEmpty()) { + android.util.Log.i("WebPhotoFetcher", "reap declined, fetch in flight") + return + } val wv = webView ?: return webView = null warmed = false diff --git a/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt b/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt index b220ad06..480653d5 100644 --- a/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt +++ b/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt @@ -74,19 +74,33 @@ class WebPopularTimesFetcher @Inject constructor( main.postDelayed(r, delayMs) } - /** Destroy the WebView immediately. Must run on the main thread (WebView requirement). */ - private fun reapNow() { + /** Destroy the WebView. Main thread only (WebView requirement). In-flight aware, both ways: + * with a fetch in flight ([pending] non-empty) a non-forced reap DECLINES - destroying the + * view kills the injected scraper and turns a live fetch into an empty result, and the fetch's + * own finally re-arms the reap moments later anyway. A FORCED reap (critical trim - the OS is + * about to start killing) destroys regardless, then drains [pending] so the stranded fetch + * fails fast as empty instead of parking in `deferred.await()` for the full + * [TOTAL_TIMEOUT_MS]. */ + private fun reapNow(force: Boolean = false) { + if (!force && pending.isNotEmpty()) { + android.util.Log.i("WebPopularTimesFetcher", "reap declined, fetch in flight") + return + } webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } webView = null warm = null // ensureWarm re-runs the warm sequence on the next fetch + pending.keys.toList().forEach { id -> pending.remove(id)?.complete("") } } init { // Under real memory pressure the 120 s idle timer is far too slow - the OS is asking for // memory NOW and a Chromium renderer is one of the largest things we hold (issue #83). - // Reap on the main thread, since WebView.destroy() requires it. + // Reap on the main thread, since WebView.destroy() requires it. Only a CRITICAL trim tears + // down mid-fetch; a merely-severe one declines while a fetch is in flight (see reapNow). app.vela.ui.MemoryPressure.register { level -> - if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { cancelReap(); reapNow() } + if (app.vela.ui.MemoryPressure.isSevere(level)) { + main.post { cancelReap(); reapNow(force = app.vela.ui.MemoryPressure.isCritical(level)) } + } } } diff --git a/app/src/main/java/app/vela/web/WebReviewsFetcher.kt b/app/src/main/java/app/vela/web/WebReviewsFetcher.kt index a2465860..1ccf9b6e 100644 --- a/app/src/main/java/app/vela/web/WebReviewsFetcher.kt +++ b/app/src/main/java/app/vela/web/WebReviewsFetcher.kt @@ -54,32 +54,54 @@ class WebReviewsFetcher @Inject constructor( @Volatile private var webView: WebView? = null private var reap: Runnable? = null + /** Run [block] on the main thread, inline when already there. All reap bookkeeping goes through + * here: [reap] is touched by the trim listener (main) and by the fetch path (the caller's + * dispatcher), and scheduling is a read-modify-write - main-confinement removes the race by + * construction (WebPhotoFetcher's fix, issue #83 follow-up). */ + private fun onMain(block: () -> Unit) { + if (Looper.myLooper() == Looper.getMainLooper()) block() else main.post(block) + } + /** Free the WebView after a quiet period (issue #182): a warm fetcher pins a full * maps.google.com page for the rest of the session, and several warm fetchers at once is * real memory. The next fetch after a reap just re-creates it. */ - private fun scheduleReap() { + private fun scheduleReap() = onMain { reap?.let(main::removeCallbacks) - val r = Runnable { reapNow() } + val r = Runnable { reap = null; reapNow() } reap = r main.postDelayed(r, REAP_IDLE_MS) } - /** Destroy the WebView immediately. Must run on the main thread (WebView requirement). */ - private fun reapNow() { + /** Destroy the WebView. Main thread only (WebView requirement). In-flight aware, both ways: + * with a fetch in flight ([pending] non-empty) a non-forced reap DECLINES - destroying the + * view kills the injected scraper and turns a live fetch into an empty panel, and the fetch's + * own finally re-arms the reap moments later anyway. A FORCED reap (critical trim - the OS is + * about to start killing) destroys regardless, then drains [pending] so the stranded fetch + * fails fast as empty instead of parking in `deferred.await()` for the full [TOTAL_TIMEOUT_MS] + * while holding [mutex] (which would stall every queued fetch behind a 45 s hang). */ + private fun reapNow(force: Boolean = false) { + if (!force && pending.isNotEmpty()) { + android.util.Log.i("WebReviewsFetcher", "reap declined, fetch in flight") + return + } webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } webView = null + pending.keys.toList().forEach { id -> pending.remove(id)?.complete("") } } init { // Under real memory pressure the 120 s idle timer is far too slow - the OS is asking for // memory NOW and a Chromium renderer is one of the largest things we hold (issue #83). - // Reap on the main thread, since WebView.destroy() requires it. + // Reap on the main thread, since WebView.destroy() requires it. Only a CRITICAL trim tears + // down mid-fetch; a merely-severe one declines while a fetch is in flight (see reapNow). app.vela.ui.MemoryPressure.register { level -> - if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { cancelReap(); reapNow() } + if (app.vela.ui.MemoryPressure.isSevere(level)) { + main.post { cancelReap(); reapNow(force = app.vela.ui.MemoryPressure.isCritical(level)) } + } } } - private fun cancelReap() { + private fun cancelReap() = onMain { reap?.let(main::removeCallbacks) reap = null } diff --git a/app/src/main/java/app/vela/web/WebStopDeparturesFetcher.kt b/app/src/main/java/app/vela/web/WebStopDeparturesFetcher.kt index 32d1dff4..af133a33 100644 --- a/app/src/main/java/app/vela/web/WebStopDeparturesFetcher.kt +++ b/app/src/main/java/app/vela/web/WebStopDeparturesFetcher.kt @@ -48,33 +48,55 @@ class WebStopDeparturesFetcher @Inject constructor( @Volatile private var webView: WebView? = null private var reap: Runnable? = null + /** Run [block] on the main thread, inline when already there. All reap bookkeeping goes through + * here: [reap] is touched by the trim listener (main) and by the fetch path (the caller's + * dispatcher), and scheduling is a read-modify-write - main-confinement removes the race by + * construction (WebPhotoFetcher's fix, issue #83 follow-up). */ + private fun onMain(block: () -> Unit) { + if (Looper.myLooper() == Looper.getMainLooper()) block() else main.post(block) + } + /** Free the WebView after a quiet period. A warm fetcher otherwise pins a full * maps.google.com page (DOM + renderer) for the rest of the session, and several warm * fetchers at once is real memory pressure (issue #182). The next fetch after a reap * just re-creates it - a one-off warm-up, only after minutes of not using the feature. */ - private fun scheduleReap() { + private fun scheduleReap() = onMain { reap?.let(main::removeCallbacks) - val r = Runnable { reapNow() } + val r = Runnable { reap = null; reapNow() } reap = r main.postDelayed(r, REAP_IDLE_MS) } - /** Destroy the WebView immediately. Must run on the main thread (WebView requirement). */ - private fun reapNow() { + /** Destroy the WebView. Main thread only (WebView requirement). In-flight aware, both ways: + * with a fetch in flight ([pending] non-empty) a non-forced reap DECLINES - destroying the + * view kills the injected scraper and turns a live fetch into an empty board, and the fetch's + * own finally re-arms the reap moments later anyway. A FORCED reap (critical trim - the OS is + * about to start killing) destroys regardless, then drains [pending] so the stranded fetch + * fails fast as empty instead of parking in `deferred.await()` for the full [TOTAL_TIMEOUT_MS] + * while holding [mutex]. */ + private fun reapNow(force: Boolean = false) { + if (!force && pending.isNotEmpty()) { + android.util.Log.i("WebStopDeparturesFetcher", "reap declined, fetch in flight") + return + } webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } webView = null + pending.keys.toList().forEach { id -> pending.remove(id)?.complete("") } } init { // Under real memory pressure the 120 s idle timer is far too slow - the OS is asking for // memory NOW and a Chromium renderer is one of the largest things we hold (issue #83). - // Reap on the main thread, since WebView.destroy() requires it. + // Reap on the main thread, since WebView.destroy() requires it. Only a CRITICAL trim tears + // down mid-fetch; a merely-severe one declines while a fetch is in flight (see reapNow). app.vela.ui.MemoryPressure.register { level -> - if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { cancelReap(); reapNow() } + if (app.vela.ui.MemoryPressure.isSevere(level)) { + main.post { cancelReap(); reapNow(force = app.vela.ui.MemoryPressure.isCritical(level)) } + } } } - private fun cancelReap() { + private fun cancelReap() = onMain { reap?.let(main::removeCallbacks) reap = null } diff --git a/app/src/main/res/values/strings.xml b/app/src/main/res/values/strings.xml index 37569db3..55632c9b 100644 --- a/app/src/main/res/values/strings.xml +++ b/app/src/main/res/values/strings.xml @@ -541,6 +541,7 @@ Every language Vela supports English, Chinese, Japanese, Korean, Cantonese English only + English only, uses the least memory Active Use Download (%1$d MB) diff --git a/app/src/test/java/app/vela/data/FlockCamerasTest.kt b/app/src/test/java/app/vela/data/FlockCamerasTest.kt new file mode 100644 index 00000000..4836beac --- /dev/null +++ b/app/src/test/java/app/vela/data/FlockCamerasTest.kt @@ -0,0 +1,108 @@ +package app.vela.data + +import java.io.ByteArrayInputStream +import java.io.ByteArrayOutputStream +import java.util.zip.GZIPOutputStream +import org.junit.After +import org.junit.Assert.assertEquals +import org.junit.Assert.assertTrue +import org.junit.Test + +/** + * The CSR cell index behind the ALPR layer, exercised through the same gzipped-TSV parse the app + * uses. The generation-swap cases exist because the index was once published as six separate + * fields, which could TEAR against a concurrent viewport scan during refresh()'s hot-swap (new + * `cellKeys` binary-searched against old `cellStart` = index out of bounds, an app crash with the + * camera layer on). A single published snapshot cannot tear; these tests pin the behaviour that + * refactor must keep: queries are correct before, between, and after swaps, and a swap to a LARGER + * dataset leaves every query in bounds. + */ +class FlockCamerasTest { + + @After fun reset() = FlockCameras.resetForTest() + + private fun gzTsv(rows: List>): ByteArrayInputStream { + val bos = ByteArrayOutputStream() + GZIPOutputStream(bos).bufferedWriter().use { w -> + rows.forEach { (la, lo, op) -> w.write("$la\t$lo\t$op\n") } + } + return ByteArrayInputStream(bos.toByteArray()) + } + + @Test + fun `unloaded queries are empty, not crashes`() { + assertTrue(FlockCameras.inBox(-90.0, -180.0, 90.0, 180.0).isEmpty()) + assertEquals(0, FlockCameras.size) + } + + @Test + fun `inBox finds exactly the cameras inside the box`() { + FlockCameras.loadFromForTest( + gzTsv( + listOf( + Triple(40.0, -75.0, "opA"), + Triple(40.05, -75.05, "opB"), + Triple(41.0, -75.0, "far-north"), + Triple(40.0, -76.0, "far-west"), + ), + ), + ) + assertEquals(4, FlockCameras.size) + val hit = FlockCameras.inBox(39.9, -75.2, 40.2, -74.9) + assertEquals(setOf("opA", "opB"), hit.map { it.operator }.toSet()) + } + + @Test + fun `cameras straddling many grid cells are all found`() { + // 0.1 deg cells: place one camera per cell across a 3x3 block and query the whole block. + val rows = ArrayList>() + for (i in 0..2) for (j in 0..2) rows.add(Triple(10.05 + i * 0.1, 20.05 + j * 0.1, "c$i$j")) + FlockCameras.loadFromForTest(gzTsv(rows)) + val hit = FlockCameras.inBox(10.0, 20.0, 10.3, 20.3) + assertEquals(9, hit.size) + } + + @Test + fun `swap to a larger generation keeps every query in bounds and correct`() { + FlockCameras.loadFromForTest(gzTsv(listOf(Triple(40.0, -75.0, "old")))) + assertEquals(1, FlockCameras.inBox(39.0, -76.0, 41.0, -74.0).size) + // The refresh() path: a bigger dataset with MORE occupied cells replaces the index. Under + // the torn-fields publish this was the crash shape (new keys, old starts); with a snapshot + // it must simply answer from the new generation. + val rows = (0 until 500).map { Triple(30.0 + it * 0.01, -100.0 + it * 0.01, "new$it") } + FlockCameras.loadFromForTest(gzTsv(rows)) + assertEquals(500, FlockCameras.size) + assertTrue(FlockCameras.inBox(39.0, -76.0, 41.0, -74.0).isEmpty()) // "old" is gone + assertEquals(500, FlockCameras.inBox(29.0, -101.0, 36.0, -94.0).size) + } + + @Test + fun `along finds cameras near the route and only those`() { + FlockCameras.loadFromForTest( + gzTsv( + listOf( + Triple(40.0005, -75.0, "on-route"), // ~55 m off the segment + Triple(40.05, -75.0, "far"), // ~5.5 km off + ), + ), + ) + val poly = listOf( + app.vela.core.model.LatLng(40.0, -75.01), + app.vela.core.model.LatLng(40.0, -74.99), + ) + assertEquals(listOf("on-route"), FlockCameras.along(poly, meters = 120.0).map { it.operator }) + } + + @Test + fun `malformed lines are skipped, not fatal`() { + val bos = ByteArrayOutputStream() + GZIPOutputStream(bos).bufferedWriter().use { w -> + w.write("not-a-number\t-75.0\topX\n") + w.write("40.0\n") + w.write("40.0\t-75.0\topGood\n") + } + FlockCameras.loadFromForTest(ByteArrayInputStream(bos.toByteArray())) + assertEquals(1, FlockCameras.size) + assertEquals("opGood", FlockCameras.inBox(39.0, -76.0, 41.0, -74.0).single().operator) + } +} diff --git a/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt b/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt index 4b73883c..31e0ef6f 100644 --- a/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt +++ b/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt @@ -6,7 +6,6 @@ import app.vela.core.config.JsTransforms import app.vela.core.diag.DiagLog import app.vela.core.data.CalibrationNeededException import app.vela.core.data.CategoryFilter -import app.vela.core.data.LowRamMode import app.vela.core.data.MapDataSource import app.vela.core.data.RouteEngine import app.vela.core.data.RouteGeometry @@ -188,8 +187,9 @@ class GoogleMapsDataSource @Inject constructor( // buffer - the response String, the stripped copy, the JsonElement DOM - is allocated INSIDE // `withPermit`. At most 4 of those exist at once no matter how many terms are queued behind // them, so going 15 -> 8 changes how many WAVES the fan-out takes, not how much is resident - // at the peak. The levers that do move the peak are the permit count and the response size, - // and the `!7i` pool halving below is the one being used. + // at the peak. The levers that do move the peak are the permit count and the response size; + // neither is currently reduced on low-RAM (the `!7i` pool halving was reverted - it cost + // ranked pins, see fetchTerm below). // // And its stated justification was false. It kept "school" and "park" on the grounds that // only they lack a second source once the ambient layer is active. In fact NOTHING has a @@ -206,10 +206,16 @@ class GoogleMapsDataSource @Inject constructor( val pb = SearchPb.build(term, center, cal.searchPb) .replaceFirst(Regex("!1d[0-9.]+"), "!1d${spanMeters.toInt()}") .replaceFirst(Regex("!4f[0-9.]+"), "!4f${String.format(java.util.Locale.US, "%.1f", zoom)}") - // Deep pool per term, so zooming in can go down the rank. Halved on low-RAM: the - // pool size drives the RESPONSE BODY size, and the body is what gets buffered - // and DOM-parsed per term (issue #83). - .replaceFirst(Regex("!7i\\d+"), if (LowRamMode.enabled) "!7i30" else "!7i60") + // Deep pool per term, so zooming in can go down the rank - the SAME depth on + // every device. This was briefly halved to !7i30 on low-RAM (issue #83, "the + // body is what gets buffered per term"), which broke the project's no-regression + // rule the quiet way: ranks 31-60 of every term simply never rendered on the + // constrained phones, and with the basemap poi_r* layers disabled wholesale + // while ambient POIs exist, those places had no second source - an absent gym or + // pharmacy, not a smaller buffer. The transient-parse peak is governed by the + // Semaphore(4) fan-out bound, not the per-term pool, so the halving saved little + // and cost pins; if parse buffers ever need shrinking, drop the PERMITS. + .replaceFirst(Regex("!7i\\d+"), "!7i60") val url = "${cal.searchEndpoint}&q=${term.enc()}&pb=${pb.enc()}".localized() SearchParser.parse(term, GoogleResponse.parse(get(url)), center, cal.paths).places }.getOrDefault(emptyList()) diff --git a/core/src/main/java/app/vela/core/voice/SpeechText.kt b/core/src/main/java/app/vela/core/voice/SpeechText.kt index ad43f769..c6e47831 100644 --- a/core/src/main/java/app/vela/core/voice/SpeechText.kt +++ b/core/src/main/java/app/vela/core/voice/SpeechText.kt @@ -124,11 +124,136 @@ object SpeechText { .trim('"', '“', '”') .trimEnd('.', '!', '?', ',', ';', ':', '…') .trim() + .let(::unshout) + .let(::spokenNumbersToDigits) private val BRACKET_TAG = Regex("\\[[^\\]]*]") private val BRACKET_TAIL = Regex("\\[[^\\]]*$") private val WHITESPACE = Regex("\\s+") + /** The librispeech-trained engines (Zipformer small) emit ALL CAPS. A search query has no use + * for shouting; mixed-case prose (Whisper) passes through untouched. */ + private fun unshout(s: String): String = + if (s.any { it.isLetter() } && s.none { it.isLowerCase() }) s.lowercase() else s + + /** + * Inverse text normalization for SEARCH transcripts: spoken number words become digits, because + * an address query must reach the geocoder as "123 main street", never "one twenty three main + * street". Whisper writes digits itself (this is a no-op on its output); the librispeech-trained + * Zipformer emits spoken-form words. Number words in any other language pass through untouched - + * the English vocabulary simply doesn't match - so this is safe app-wide. + * + * Spoken addresses use JUXTAPOSITION, not place value: "one twenty three" is 1|23 -> "123", + * "twelve thirty four" is 12|34 -> "1234", "one oh five" is 1|0|5 -> "105". So each spoken + * GROUP converts on its own and adjacent groups concatenate; within a group the normal rules + * apply ("twenty three" -> 23, "three hundred" -> 300, "five thousand two hundred" -> 5200). + * A trailing ordinal closes the run with its suffix ("one hundred twenty fifth" -> "125th", + * for numbered streets). "oh" counts as a zero only INSIDE a run, so the interjection alone is + * never touched. + */ + fun spokenNumbersToDigits(s: String): String { + val words = s.split(' ') + val out = StringBuilder() + var i = 0 + while (i < words.size) { + val run = parseNumberRun(words, i) + if (run == null) { + if (out.isNotEmpty()) out.append(' ') + out.append(words[i]); i++ + } else { + if (out.isNotEmpty()) out.append(' ') + out.append(run.first); i = run.second + } + } + return out.toString() + } + + /** Parse the longest spoken-number run starting at [start]; null if [start] isn't a number + * word. Returns the rendered digits (with ordinal suffix if the run ends on one) and the + * index PAST the run. */ + @Suppress("ReturnCount", "CyclomaticComplexMethod", "LongMethod") + private fun parseNumberRun(words: List, start: Int): Pair? { + val groups = ArrayList() + // A group is total + current: "thousand" banks (current * 1000) into total, "hundred" + // multiplies current, tens/teens/units build current. Group value = total + current. + var total = 0L + var current = 0L + var started = false // distinguishes "no group" from a group currently worth 0 + var canAddTens = false // after "hundred"/"thousand" a tens/teen ADDS instead of juxtaposing + var canAddUnit = false // after a tens word ("twenty"), a unit ADDS ("three" -> 23) + var ordinalSuffix: String? = null + var i = start + fun closeGroup() { + if (started) { groups.add(total + current); total = 0; current = 0; started = false } + canAddTens = false; canAddUnit = false + } + loop@ while (i < words.size && ordinalSuffix == null) { + val w = words[i].lowercase().trimEnd(',') + // Hyphenated compounds ("twenty-three") arrive as one token: handle the parts in turn. + val parts = if ('-' in w) w.split('-') else listOf(w) + for (p in parts) { + val unit = CARD.indexOf(p) // one..nine -> 1..9 (index 0 is "") + val teen = TEEN.indexOf(p) // ten..nineteen -> 0..9 + val tens = TENS_CARD.indexOf(p) // twenty..ninety -> 2..9 (0,1 unused) + val ordU = ORD1.indexOf(p) // first..ninth + val ordTeen = TEEN_ORD.indexOf(p) // tenth..nineteenth + val ordTens = TENS_ORD.indexOf(p) // twentieth..ninetieth + when { + p == "zero" || (p == "oh" && (started || groups.isNotEmpty())) -> { + closeGroup(); groups.add(0) + } + unit > 0 -> { + if (started && current > 0 && !canAddUnit && !canAddTens) closeGroup() // juxtaposed + current += unit; started = true; canAddUnit = false; canAddTens = false + } + teen >= 0 -> { + if (started && current > 0 && !canAddTens) closeGroup() + current += 10 + teen; started = true; canAddTens = false; canAddUnit = false + } + tens >= 2 -> { + if (started && current > 0 && !canAddTens) closeGroup() + current += tens * 10; started = true; canAddTens = false; canAddUnit = true + } + p == "hundred" && current in 1..99 -> { + current *= 100; canAddTens = true; canAddUnit = false + } + p == "thousand" && current in 1..999 -> { + total += current * 1000; current = 0; canAddTens = true; canAddUnit = false + } + ordU > 0 || ordTeen >= 0 || ordTens >= 2 -> { + val v = when { + ordU > 0 -> ordU.toLong() + ordTeen >= 0 -> (10 + ordTeen).toLong() + else -> ordTens * 10L + } + // A LONE ordinal converts only for tenth+: bare "first".."ninth" are common + // non-numeric English ("second opinion clinic") and rewriting them is wrong + // more often than right, while "thirteenth"/"fortieth" in a query is a + // numbered street. Attached to a number ("forty second") it always converts. + if (!started && groups.isEmpty() && ordU > 0) return null + if (started && current > 0 && !canAddUnit && !canAddTens) closeGroup() + current += v; started = true + ordinalSuffix = ordSuffix(current) + } + else -> break@loop // not a number word: the run ends before this token + } + } + i++ + } + closeGroup() + if (groups.isEmpty()) return null + val digits = groups.joinToString("") { it.toString() } + return (digits + (ordinalSuffix ?: "")) to i + } + + private fun ordSuffix(v: Long): String = when { + v % 100 in 11..13 -> "th" + v % 10 == 1L -> "st" + v % 10 == 2L -> "nd" + v % 10 == 3L -> "rd" + else -> "th" + } + private fun twoDigitOrdinal(r: Int): String = when { r in 10..19 -> TEEN_ORD[r - 10] r % 10 == 0 -> TENS_ORD[r / 10] @@ -141,6 +266,10 @@ object SpeechText { private val STREET_ORDINAL = Regex("\\b([1-9])(\\d\\d)(?:st|nd|rd|th)\\b") private val CARD = arrayOf("", "one", "two", "three", "four", "five", "six", "seven", "eight", "nine") + private val TEEN = arrayOf( + "ten", "eleven", "twelve", "thirteen", "fourteen", + "fifteen", "sixteen", "seventeen", "eighteen", "nineteen", + ) private val ORD1 = arrayOf("", "first", "second", "third", "fourth", "fifth", "sixth", "seventh", "eighth", "ninth") private val TEEN_ORD = arrayOf( "tenth", "eleventh", "twelfth", "thirteenth", "fourteenth", diff --git a/core/src/test/java/app/vela/core/voice/SpeechTextTest.kt b/core/src/test/java/app/vela/core/voice/SpeechTextTest.kt index 0e7f9a98..4512162e 100644 --- a/core/src/test/java/app/vela/core/voice/SpeechTextTest.kt +++ b/core/src/test/java/app/vela/core/voice/SpeechTextTest.kt @@ -204,4 +204,55 @@ class SpeechTextTest { assertEquals("St. Paul", clean("St. Paul")) assertEquals("J.C. Penney", clean("J.C. Penney.")) } + + // ---- inverse text normalization (spoken numbers -> digits; Zipformer emits word form) ---- + + @Test fun `all-caps librispeech output is lowercased, mixed case is not`() { + assertEquals("coffee shops near me", clean("COFFEE SHOPS NEAR ME")) + assertEquals("Coffee near St. Paul", clean("Coffee near St. Paul")) + } + + @Test fun `spoken address numbers become digits by juxtaposition`() { + assertEquals("123 main street", clean("ONE TWENTY THREE MAIN STREET")) + assertEquals("1234 elm avenue", clean("twelve thirty four elm avenue")) + assertEquals("105 broad street", clean("one oh five broad street")) + assertEquals("90 west road", clean("ninety west road")) + assertEquals("6000 south street", clean("six thousand south street")) + } + + @Test fun `place-value groups combine before juxtaposition`() { + assertEquals("300 park avenue", clean("three hundred park avenue")) + assertEquals("125 court", clean("one hundred twenty five court")) + assertEquals("5200 ridge line", clean("five thousand two hundred ridge line")) + assertEquals("23", clean("twenty-three")) + } + + @Test fun `numbered streets keep their ordinal suffix`() { + assertEquals("125th street", clean("ONE HUNDRED TWENTY FIFTH STREET")) + assertEquals("42nd street", clean("forty second street")) + assertEquals("13th avenue", clean("thirteenth avenue")) + assertEquals("21st street", clean("twenty first street")) + } + + @Test fun `lone ordinal words stay words`() { + // "second opinion clinic" must not become "2nd opinion clinic"; bare ordinals are common + // non-numeric English. ("first avenue" also geocodes fine as written.) + assertEquals("second opinion clinic", clean("second opinion clinic")) + assertEquals("first avenue", clean("first avenue")) + } + + @Test fun `oh is a zero only inside a number run`() { + assertEquals("oh coffee", clean("oh coffee")) + assertEquals("102 main", clean("one oh two main")) + } + + @Test fun `digits from Whisper pass through unchanged`() { + assertEquals("123 Main Street", clean("123 Main Street.")) + assertEquals("42nd Street", clean("42nd Street")) + } + + @Test fun `non-English number words are untouched`() { + assertEquals("uno dos tres calle mayor", clean("uno dos tres calle mayor")) + assertEquals("einhundert dreiundzwanzig", clean("einhundert dreiundzwanzig")) + } } From ffed0ad749e8dc99188921c87436ba60c2a4fa60 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Thu, 23 Jul 2026 19:46:45 -0400 Subject: [PATCH 19/21] Translate the Zipformer engine caption into all 14 locales The translation-completeness gate caught settings_asr_langs_zipformer existing only in English; each locale follows its own moonshine-caption phrasing. --- app/src/main/res/values-de/strings.xml | 1 + app/src/main/res/values-es/strings.xml | 1 + app/src/main/res/values-fr/strings.xml | 1 + app/src/main/res/values-it/strings.xml | 1 + app/src/main/res/values-iw/strings.xml | 1 + app/src/main/res/values-ja/strings.xml | 1 + app/src/main/res/values-nl/strings.xml | 1 + app/src/main/res/values-pl/strings.xml | 1 + app/src/main/res/values-pt/strings.xml | 1 + app/src/main/res/values-ru/strings.xml | 1 + app/src/main/res/values-sv/strings.xml | 1 + app/src/main/res/values-uk/strings.xml | 1 + app/src/main/res/values-zh-rTW/strings.xml | 1 + app/src/main/res/values-zh/strings.xml | 1 + 14 files changed, 14 insertions(+) diff --git a/app/src/main/res/values-de/strings.xml b/app/src/main/res/values-de/strings.xml index 3f21e7e4..44110d02 100644 --- a/app/src/main/res/values-de/strings.xml +++ b/app/src/main/res/values-de/strings.xml @@ -489,6 +489,7 @@ Alle von Vela unterstützten Sprachen Englisch, Chinesisch, Japanisch, Koreanisch, Kantonesisch Nur Englisch + Nur Englisch, benötigt am wenigsten Speicher Aktiv Verwenden Herunterladen (%1$d MB) diff --git a/app/src/main/res/values-es/strings.xml b/app/src/main/res/values-es/strings.xml index f5de8fdd..4caf745d 100644 --- a/app/src/main/res/values-es/strings.xml +++ b/app/src/main/res/values-es/strings.xml @@ -489,6 +489,7 @@ Todos los idiomas que admite Vela Inglés, chino, japonés, coreano, cantonés Solo inglés + Solo inglés, usa menos memoria Activo Usar Descargar (%1$d MB) diff --git a/app/src/main/res/values-fr/strings.xml b/app/src/main/res/values-fr/strings.xml index 152ad4b3..eb4a1942 100644 --- a/app/src/main/res/values-fr/strings.xml +++ b/app/src/main/res/values-fr/strings.xml @@ -489,6 +489,7 @@ Toutes les langues prises en charge par Vela Anglais, chinois, japonais, coréen, cantonais Anglais uniquement + Anglais uniquement, utilise le moins de mémoire Actif Utiliser Télécharger (%1$d Mo) diff --git a/app/src/main/res/values-it/strings.xml b/app/src/main/res/values-it/strings.xml index 0dab89a6..cfd4f901 100644 --- a/app/src/main/res/values-it/strings.xml +++ b/app/src/main/res/values-it/strings.xml @@ -489,6 +489,7 @@ Tutte le lingue supportate da Vela Inglese, cinese, giapponese, coreano, cantonese Solo inglese + Solo inglese, usa meno memoria Attivo Usa Scarica (%1$d MB) diff --git a/app/src/main/res/values-iw/strings.xml b/app/src/main/res/values-iw/strings.xml index 9169ce5e..33f214c8 100644 --- a/app/src/main/res/values-iw/strings.xml +++ b/app/src/main/res/values-iw/strings.xml @@ -476,6 +476,7 @@ כל השפות ש-Vela תומכת בהן אנגלית, סינית, יפנית, קוריאנית, קנטונזית אנגלית בלבד + אנגלית בלבד, צורכת הכי מעט זיכרון פעיל השתמש הורד (%1$d מ\"ב) diff --git a/app/src/main/res/values-ja/strings.xml b/app/src/main/res/values-ja/strings.xml index 78a08dbf..971b4f84 100644 --- a/app/src/main/res/values-ja/strings.xml +++ b/app/src/main/res/values-ja/strings.xml @@ -496,6 +496,7 @@ Vela が対応するすべての言語 英語、中国語、日本語、韓国語、広東語 英語のみ + 英語のみ、メモリ使用量が最小 使用中 使う ダウンロード(%1$d MB) diff --git a/app/src/main/res/values-nl/strings.xml b/app/src/main/res/values-nl/strings.xml index e32a6e8c..c0c2865c 100644 --- a/app/src/main/res/values-nl/strings.xml +++ b/app/src/main/res/values-nl/strings.xml @@ -489,6 +489,7 @@ Alle talen die Vela ondersteunt Engels, Chinees, Japans, Koreaans, Kantonees Alleen Engels + Alleen Engels, gebruikt het minste geheugen Actief Gebruiken Downloaden (%1$d MB) diff --git a/app/src/main/res/values-pl/strings.xml b/app/src/main/res/values-pl/strings.xml index 38f6b230..28619191 100644 --- a/app/src/main/res/values-pl/strings.xml +++ b/app/src/main/res/values-pl/strings.xml @@ -495,6 +495,7 @@ Wszystkie języki obsługiwane przez Vela angielski, chiński, japoński, koreański, kantoński Tylko angielski + Tylko angielski, zużywa najmniej pamięci Aktywny Użyj Pobierz (%1$d MB) diff --git a/app/src/main/res/values-pt/strings.xml b/app/src/main/res/values-pt/strings.xml index 3916c288..139957bc 100644 --- a/app/src/main/res/values-pt/strings.xml +++ b/app/src/main/res/values-pt/strings.xml @@ -489,6 +489,7 @@ Todos os idiomas suportados pelo Vela Inglês, chinês, japonês, coreano, cantonês Apenas inglês + Apenas inglês, usa menos memória Ativo Usar Transferir (%1$d MB) diff --git a/app/src/main/res/values-ru/strings.xml b/app/src/main/res/values-ru/strings.xml index 46dc17c8..140c9393 100644 --- a/app/src/main/res/values-ru/strings.xml +++ b/app/src/main/res/values-ru/strings.xml @@ -495,6 +495,7 @@ Все языки, которые поддерживает Vela английский, китайский, японский, корейский, кантонский Только английский + Только английский, использует меньше всего памяти Активно Выбрать Загрузить (%1$d МБ) diff --git a/app/src/main/res/values-sv/strings.xml b/app/src/main/res/values-sv/strings.xml index 9046109f..d289440a 100644 --- a/app/src/main/res/values-sv/strings.xml +++ b/app/src/main/res/values-sv/strings.xml @@ -489,6 +489,7 @@ Alla språk som Vela stöder Engelska, kinesiska, japanska, koreanska, kantonesiska Endast engelska + Endast engelska, använder minst minne Aktiv Använd Ladda ner (%1$d MB) diff --git a/app/src/main/res/values-uk/strings.xml b/app/src/main/res/values-uk/strings.xml index 922e030c..1bd1837f 100644 --- a/app/src/main/res/values-uk/strings.xml +++ b/app/src/main/res/values-uk/strings.xml @@ -495,6 +495,7 @@ Усі мови, які підтримує Vela англійська, китайська, японська, корейська, кантонська Лише англійська + Лише англійська, використовує найменше пам'яті Активно Використати Завантажити (%1$d МБ) diff --git a/app/src/main/res/values-zh-rTW/strings.xml b/app/src/main/res/values-zh-rTW/strings.xml index 60c36038..2d1ac73d 100644 --- a/app/src/main/res/values-zh-rTW/strings.xml +++ b/app/src/main/res/values-zh-rTW/strings.xml @@ -496,6 +496,7 @@ Vela 支援的所有語言 英語、中文、日語、韓語、粤語 僅英語 + 僅英語,佔用記憶體最少 使用中 使用 下載(%1$d MB) diff --git a/app/src/main/res/values-zh/strings.xml b/app/src/main/res/values-zh/strings.xml index 2e3ddba8..cae4b482 100644 --- a/app/src/main/res/values-zh/strings.xml +++ b/app/src/main/res/values-zh/strings.xml @@ -496,6 +496,7 @@ Vela 支持的所有语言 英语、汉语、日语、韩语、粤语 仅英语 + 仅英语,占用内存最少 使用中 使用 下载(%1$d MB) From c4c06f891e855f9a64f32ef796b7403b1be06bc0 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Thu, 23 Jul 2026 19:50:29 -0400 Subject: [PATCH 20/21] Remove the Zipformer engine; keep the transcript normalization Resident size was never the only bar: Zipformer small is a 2023 librispeech (audiobook-domain) model, and a maps app lives on the proper nouns that domain mishears - no post-processing fixes that. With the idle reaper making Whisper's ~214 MB hold transient rather than session-long, the low-memory engine stopped paying for its accuracy cost. The engine entry, config branch, picker caption and its 14 translations go; the measurement lessons (NeMo ballooning, Moonshine's four sessions, the reap-delta isolation method) stay in AGENTS.md so the ground isn't re-trodden. spokenNumbersToDigits + the ALL-CAPS unshout stay in cleanSearchTranscript with their tests: they are a no-op on the three digit-writing engines and insurance for any future word-form one. --- AGENTS.md | 32 ++++++++++--------- .../ui/settings/sections/SearchSettings.kt | 1 - app/src/main/java/app/vela/voice/AsrEngine.kt | 30 ++++++----------- .../java/app/vela/voice/WhisperRecognizer.kt | 12 ------- app/src/main/res/values-de/strings.xml | 1 - app/src/main/res/values-es/strings.xml | 1 - app/src/main/res/values-fr/strings.xml | 1 - app/src/main/res/values-it/strings.xml | 1 - app/src/main/res/values-iw/strings.xml | 1 - app/src/main/res/values-ja/strings.xml | 1 - app/src/main/res/values-nl/strings.xml | 1 - app/src/main/res/values-pl/strings.xml | 1 - app/src/main/res/values-pt/strings.xml | 1 - app/src/main/res/values-ru/strings.xml | 1 - app/src/main/res/values-sv/strings.xml | 1 - app/src/main/res/values-uk/strings.xml | 1 - app/src/main/res/values-zh-rTW/strings.xml | 1 - app/src/main/res/values-zh/strings.xml | 1 - app/src/main/res/values/strings.xml | 1 - 19 files changed, 26 insertions(+), 64 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 320897da..59bcd6d4 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -845,13 +845,15 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): Conformer CTC small is a 46 MB int8 file that ballooned to ~760 MB-1.2 GB PSS through onnxruntime on the M5 (rejected); Moonshine's small weights still cost ~212 MB across its four ORT sessions - no lighter resident than Whisper tiny's ~214 MB (both same-protocol launch - deltas, 32-bit M5). Zipformer small (26 MB encoder, one tiny decoder/joiner pair) is the - structural low-memory candidate; its isolated reap-delta number is tracked in PR #86 and MUST - be written here before any low-RAM default is switched to it. "Encoder-only means small" and - "fewer MB means less RAM" were both device-refuted in one afternoon. The clean way to isolate - one model's resident cost: let the IDLE REAP fire (it releases ONLY the recognizer) and diff - `Native Heap` pre/post inside one settled process - launch-to-launch totals swing +-150 MB with - map content and prove nothing at engine granularity. + deltas, 32-bit M5). k2 Zipformer small (26 MB encoder) was built, wired and then REMOVED + before merge: resident size is not the only bar - it is a 2023 librispeech (audiobook-domain) + model, and a maps app lives on the proper nouns that domain mishears; no post-processing fixes + that. The 267 MB Whisper hold is TRANSIENT now (idle reap), which is what actually makes it + viable on small phones. "Encoder-only means small" and "fewer MB means less RAM" were both + device-refuted in one afternoon. The clean way to isolate one model's resident cost: let the + IDLE REAP fire (it releases ONLY the recognizer) and diff `Native Heap` pre/post inside one + settled process - launch-to-launch totals swing +-150 MB with map content and prove nothing at + engine granularity. - **Do not reach for a device gate when an IDLE gate will do.** The warm-at-startup behaviour was first made low-RAM-conditional, which protected the instant-first-mic-tap UX on roomier phones but left them holding 267 MB all session. Reaping on idle keeps that UX AND reclaims the memory @@ -1743,14 +1745,14 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): D-pad-focusable `Dialog`, Done auto-focuses); wiring + the RECORD_AUDIO launcher + the download-offer are in `MapScreen`; the Settings -> Search per-engine picker is in `SettingsScreen`. Needs `RECORD_AUDIO` (manifest; asked at the mic tap). - - **PICKABLE ENGINES via `voice/AsrEngine` (an enum catalog; first three ported upstream - 5d2a6636 / 118e7e8c / 137beea9).** Four: `WHISPER_TINY` (multilingual, ~58 MB, the `DEFAULT`), - `SENSE_VOICE` (en/zh/ja/ko/yue, ~154 MB), `MOONSHINE` (English-only, ~101 MB), - `ZIPFORMER_SMALL` (English-only, ~28 MB, this fork's low-memory pick - and it emits ALL-CAPS - spoken-form text, so `SpeechText.cleanSearchTranscript` lowercases and runs - `spokenNumbersToDigits`, the unit-tested inverse text normalization that turns "ONE TWENTY - THREE MAIN STREET" into "123 main street"; Whisper writes digits itself and passes through - unchanged). Each is an OPTIONAL download to + - **PICKABLE ENGINES via `voice/AsrEngine` (an enum catalog; ported upstream 5d2a6636 / + 118e7e8c / 137beea9).** Three: `WHISPER_TINY` (multilingual, ~58 MB, the `DEFAULT`), + `SENSE_VOICE` (en/zh/ja/ko/yue, ~154 MB), `MOONSHINE` (English-only, ~101 MB). Transcripts + pass through `SpeechText.cleanSearchTranscript`, which also lowercases ALL-CAPS output and + runs `spokenNumbersToDigits` - unit-tested inverse text normalization ("ONE TWENTY THREE + MAIN STREET" -> "123 main street") kept as insurance for any future word-form engine; the + current three write digits themselves and pass through unchanged. Each is an OPTIONAL + download to `filesDir/asr//`; `active()` is the picked engine (pref `asr_engine`), `forRecognition(lang)` falls back to Whisper when the pick can't do the app language. **Whisper stays the default and the ONLY thing onboarding / the map mic offer install** (`downloadAsrModel()` = diff --git a/app/src/main/java/app/vela/ui/settings/sections/SearchSettings.kt b/app/src/main/java/app/vela/ui/settings/sections/SearchSettings.kt index e4d5d7d8..aa4727fd 100644 --- a/app/src/main/java/app/vela/ui/settings/sections/SearchSettings.kt +++ b/app/src/main/java/app/vela/ui/settings/sections/SearchSettings.kt @@ -80,7 +80,6 @@ internal fun SearchSettingsScreen(vm: MapViewModel, onBack: () -> Unit) { app.vela.voice.AsrEngine.WHISPER_TINY -> R.string.settings_asr_langs_whisper app.vela.voice.AsrEngine.SENSE_VOICE -> R.string.settings_asr_langs_sensevoice app.vela.voice.AsrEngine.MOONSHINE -> R.string.settings_asr_langs_moonshine - app.vela.voice.AsrEngine.ZIPFORMER_SMALL -> R.string.settings_asr_langs_zipformer }, ) val meta = stringResource(R.string.settings_asr_engine_meta, langs, engine.sizeMb) diff --git a/app/src/main/java/app/vela/voice/AsrEngine.kt b/app/src/main/java/app/vela/voice/AsrEngine.kt index e7afa894..d1afff44 100644 --- a/app/src/main/java/app/vela/voice/AsrEngine.kt +++ b/app/src/main/java/app/vela/voice/AsrEngine.kt @@ -17,26 +17,23 @@ private const val VAD_FILE = "silero_vad.onnx" * no account or third-party voice app is needed. Each engine is an OPTIONAL one-time download hosted * on this repo's `asr-models` GitHub release, extracted to `filesDir/asr//`. * - * Four engines, because they trade off differently and the user picks (the first three ported from - * upstream PimpinPumpkin/Vela 5d2a6636 + 118e7e8c sizes + 137beea9 language fallback): + * Three engines, because they trade off differently and the user picks (ported from upstream + * PimpinPumpkin/Vela 5d2a6636 + 118e7e8c sizes + 137beea9 language fallback): * - [WHISPER_TINY] - the multilingual default. 99-language Whisper tiny (int8); covers every * language Vela's UI supports (incl. Hebrew, Russian, Spanish). The safe all-rounder and the * smallest download - so it stays the default and the ONLY thing the one-tap onboarding/map - * offer installs. NOT the smallest loaded: ~214 MB PSS resident (measured, 32-bit M5). + * offer installs. NOT the smallest loaded: ~214 MB PSS resident (measured, 32-bit M5) - but + * the idle reaper makes that transient, not session-long. * - [SENSE_VOICE] - FunAudioLLM SenseVoice. More accurate + faster than Whisper tiny, but only for * English, Chinese, Cantonese, Japanese, Korean. Bigger (opt-in). * - [MOONSHINE] - Useful Sensors Moonshine tiny. Lowest latency, ENGLISH ONLY. Bigger (opt-in), * and despite the small weights its four ORT sessions cost ~212 MB PSS loaded - no lighter * resident than Whisper (measured, 32-bit M5). - * - [ZIPFORMER_SMALL] - k2 Zipformer small transducer, ENGLISH ONLY. The low-memory candidate for - * the RAM-constrained feature phones this fork exists for (this fork's addition): 26 MB int8 - * encoder + kilobyte-scale decoder/joiner, no 30 s fixed context and no KV cache. Transducer - * over CTC deliberately - sherpa-onnx hotword/contextual biasing (street/place names) works on - * transducers, and the zoo has no offline English CTC small anyway. NeMo Conformer CTC small - * was tried first for this slot and REJECTED on measurement: its int8 graph ballooned to - * ~760 MB-1.2 GB PSS through onnxruntime on the M5, worse than Whisper - "encoder-only means - * small" did not survive a device number, so do not trust this entry's resident cost until the - * isolated reap-delta measurement in PR #86 is written back here. + * + * Before adding a "low-memory" fourth engine, read the measurement notes in AGENTS.md: NeMo + * Conformer CTC small (46 MB file) ballooned to ~760 MB-1.2 GB resident through onnxruntime, and + * k2 Zipformer small was built, measured and then REMOVED - librispeech-domain models mishear the + * proper nouns a maps app lives on, and no amount of post-processing fixes that. * * Whisper stays the default so no language silently regresses; the other two are opt-in via the * voice-search engine picker in Settings. This holds only metadata + a cheap install check + the @@ -82,14 +79,6 @@ enum class AsrEngine( "cached_decode.int8.onnx", "tokens.txt", VAD_FILE, ), ), - ZIPFORMER_SMALL( - id = "zipformer-small-en", - displayName = "Zipformer small", - modelType = "transducer", - sizeMb = 28, - url = "$ASR_BASE/vela-asr-zipformer-small-en.tar.bz2", - files = listOf("encoder.int8.onnx", "decoder.int8.onnx", "joiner.int8.onnx", "tokens.txt", VAD_FILE), - ), ; /** `filesDir/asr//` - the extracted archive's single top-level folder. */ @@ -109,7 +98,6 @@ enum class AsrEngine( WHISPER_TINY -> true SENSE_VOICE -> lang in SENSE_VOICE_LANGS MOONSHINE -> lang == "en" - ZIPFORMER_SMALL -> lang == "en" } companion object { diff --git a/app/src/main/java/app/vela/voice/WhisperRecognizer.kt b/app/src/main/java/app/vela/voice/WhisperRecognizer.kt index b58f2f2a..1ac5440f 100644 --- a/app/src/main/java/app/vela/voice/WhisperRecognizer.kt +++ b/app/src/main/java/app/vela/voice/WhisperRecognizer.kt @@ -14,7 +14,6 @@ import timber.log.Timber import com.k2fsa.sherpa.onnx.FeatureConfig import com.k2fsa.sherpa.onnx.OfflineModelConfig import com.k2fsa.sherpa.onnx.OfflineMoonshineModelConfig -import com.k2fsa.sherpa.onnx.OfflineTransducerModelConfig import com.k2fsa.sherpa.onnx.OfflineSenseVoiceModelConfig import com.k2fsa.sherpa.onnx.OfflineRecognizer import com.k2fsa.sherpa.onnx.OfflineRecognizerConfig @@ -285,7 +284,6 @@ class WhisperRecognizer @Inject constructor( AsrEngine.WHISPER_TINY -> l.takeIf { it in app.vela.ui.AppLocale.SUPPORTED } ?: "" AsrEngine.SENSE_VOICE -> l.takeIf { it in AsrEngine.SENSE_VOICE_LANGS } ?: "auto" AsrEngine.MOONSHINE -> "" - AsrEngine.ZIPFORMER_SMALL -> "" // English-only transducer, takes no language } } @@ -408,16 +406,6 @@ class WhisperRecognizer @Inject constructor( numThreads = 2, modelType = engine.modelType, ) - AsrEngine.ZIPFORMER_SMALL -> OfflineModelConfig( - transducer = OfflineTransducerModelConfig( - encoder = p("encoder.int8.onnx"), - decoder = p("decoder.int8.onnx"), - joiner = p("joiner.int8.onnx"), - ), - tokens = p("tokens.txt"), - numThreads = 2, - modelType = engine.modelType, - ) } val r = runCatching { OfflineRecognizer( diff --git a/app/src/main/res/values-de/strings.xml b/app/src/main/res/values-de/strings.xml index 44110d02..3f21e7e4 100644 --- a/app/src/main/res/values-de/strings.xml +++ b/app/src/main/res/values-de/strings.xml @@ -489,7 +489,6 @@ Alle von Vela unterstützten Sprachen Englisch, Chinesisch, Japanisch, Koreanisch, Kantonesisch Nur Englisch - Nur Englisch, benötigt am wenigsten Speicher Aktiv Verwenden Herunterladen (%1$d MB) diff --git a/app/src/main/res/values-es/strings.xml b/app/src/main/res/values-es/strings.xml index 4caf745d..f5de8fdd 100644 --- a/app/src/main/res/values-es/strings.xml +++ b/app/src/main/res/values-es/strings.xml @@ -489,7 +489,6 @@ Todos los idiomas que admite Vela Inglés, chino, japonés, coreano, cantonés Solo inglés - Solo inglés, usa menos memoria Activo Usar Descargar (%1$d MB) diff --git a/app/src/main/res/values-fr/strings.xml b/app/src/main/res/values-fr/strings.xml index eb4a1942..152ad4b3 100644 --- a/app/src/main/res/values-fr/strings.xml +++ b/app/src/main/res/values-fr/strings.xml @@ -489,7 +489,6 @@ Toutes les langues prises en charge par Vela Anglais, chinois, japonais, coréen, cantonais Anglais uniquement - Anglais uniquement, utilise le moins de mémoire Actif Utiliser Télécharger (%1$d Mo) diff --git a/app/src/main/res/values-it/strings.xml b/app/src/main/res/values-it/strings.xml index cfd4f901..0dab89a6 100644 --- a/app/src/main/res/values-it/strings.xml +++ b/app/src/main/res/values-it/strings.xml @@ -489,7 +489,6 @@ Tutte le lingue supportate da Vela Inglese, cinese, giapponese, coreano, cantonese Solo inglese - Solo inglese, usa meno memoria Attivo Usa Scarica (%1$d MB) diff --git a/app/src/main/res/values-iw/strings.xml b/app/src/main/res/values-iw/strings.xml index 33f214c8..9169ce5e 100644 --- a/app/src/main/res/values-iw/strings.xml +++ b/app/src/main/res/values-iw/strings.xml @@ -476,7 +476,6 @@ כל השפות ש-Vela תומכת בהן אנגלית, סינית, יפנית, קוריאנית, קנטונזית אנגלית בלבד - אנגלית בלבד, צורכת הכי מעט זיכרון פעיל השתמש הורד (%1$d מ\"ב) diff --git a/app/src/main/res/values-ja/strings.xml b/app/src/main/res/values-ja/strings.xml index 971b4f84..78a08dbf 100644 --- a/app/src/main/res/values-ja/strings.xml +++ b/app/src/main/res/values-ja/strings.xml @@ -496,7 +496,6 @@ Vela が対応するすべての言語 英語、中国語、日本語、韓国語、広東語 英語のみ - 英語のみ、メモリ使用量が最小 使用中 使う ダウンロード(%1$d MB) diff --git a/app/src/main/res/values-nl/strings.xml b/app/src/main/res/values-nl/strings.xml index c0c2865c..e32a6e8c 100644 --- a/app/src/main/res/values-nl/strings.xml +++ b/app/src/main/res/values-nl/strings.xml @@ -489,7 +489,6 @@ Alle talen die Vela ondersteunt Engels, Chinees, Japans, Koreaans, Kantonees Alleen Engels - Alleen Engels, gebruikt het minste geheugen Actief Gebruiken Downloaden (%1$d MB) diff --git a/app/src/main/res/values-pl/strings.xml b/app/src/main/res/values-pl/strings.xml index 28619191..38f6b230 100644 --- a/app/src/main/res/values-pl/strings.xml +++ b/app/src/main/res/values-pl/strings.xml @@ -495,7 +495,6 @@ Wszystkie języki obsługiwane przez Vela angielski, chiński, japoński, koreański, kantoński Tylko angielski - Tylko angielski, zużywa najmniej pamięci Aktywny Użyj Pobierz (%1$d MB) diff --git a/app/src/main/res/values-pt/strings.xml b/app/src/main/res/values-pt/strings.xml index 139957bc..3916c288 100644 --- a/app/src/main/res/values-pt/strings.xml +++ b/app/src/main/res/values-pt/strings.xml @@ -489,7 +489,6 @@ Todos os idiomas suportados pelo Vela Inglês, chinês, japonês, coreano, cantonês Apenas inglês - Apenas inglês, usa menos memória Ativo Usar Transferir (%1$d MB) diff --git a/app/src/main/res/values-ru/strings.xml b/app/src/main/res/values-ru/strings.xml index 140c9393..46dc17c8 100644 --- a/app/src/main/res/values-ru/strings.xml +++ b/app/src/main/res/values-ru/strings.xml @@ -495,7 +495,6 @@ Все языки, которые поддерживает Vela английский, китайский, японский, корейский, кантонский Только английский - Только английский, использует меньше всего памяти Активно Выбрать Загрузить (%1$d МБ) diff --git a/app/src/main/res/values-sv/strings.xml b/app/src/main/res/values-sv/strings.xml index d289440a..9046109f 100644 --- a/app/src/main/res/values-sv/strings.xml +++ b/app/src/main/res/values-sv/strings.xml @@ -489,7 +489,6 @@ Alla språk som Vela stöder Engelska, kinesiska, japanska, koreanska, kantonesiska Endast engelska - Endast engelska, använder minst minne Aktiv Använd Ladda ner (%1$d MB) diff --git a/app/src/main/res/values-uk/strings.xml b/app/src/main/res/values-uk/strings.xml index 1bd1837f..922e030c 100644 --- a/app/src/main/res/values-uk/strings.xml +++ b/app/src/main/res/values-uk/strings.xml @@ -495,7 +495,6 @@ Усі мови, які підтримує Vela англійська, китайська, японська, корейська, кантонська Лише англійська - Лише англійська, використовує найменше пам'яті Активно Використати Завантажити (%1$d МБ) diff --git a/app/src/main/res/values-zh-rTW/strings.xml b/app/src/main/res/values-zh-rTW/strings.xml index 2d1ac73d..60c36038 100644 --- a/app/src/main/res/values-zh-rTW/strings.xml +++ b/app/src/main/res/values-zh-rTW/strings.xml @@ -496,7 +496,6 @@ Vela 支援的所有語言 英語、中文、日語、韓語、粤語 僅英語 - 僅英語,佔用記憶體最少 使用中 使用 下載(%1$d MB) diff --git a/app/src/main/res/values-zh/strings.xml b/app/src/main/res/values-zh/strings.xml index cae4b482..2e3ddba8 100644 --- a/app/src/main/res/values-zh/strings.xml +++ b/app/src/main/res/values-zh/strings.xml @@ -496,7 +496,6 @@ Vela 支持的所有语言 英语、汉语、日语、韩语、粤语 仅英语 - 仅英语,占用内存最少 使用中 使用 下载(%1$d MB) diff --git a/app/src/main/res/values/strings.xml b/app/src/main/res/values/strings.xml index 55632c9b..37569db3 100644 --- a/app/src/main/res/values/strings.xml +++ b/app/src/main/res/values/strings.xml @@ -541,7 +541,6 @@ Every language Vela supports English, Chinese, Japanese, Korean, Cantonese English only - English only, uses the least memory Active Use Download (%1$d MB) From 21b08d7a0b1c6c1a7c6de1aac068d30b2f18b2b9 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Thu, 23 Jul 2026 19:55:01 -0400 Subject: [PATCH 21/21] CI: size the AAR truncation guard for 1.13.4 The 1.13.4 AAR is 48.8 MB (1.13.3 was 57 MB) and the >50 MB check failed the successfully-downloaded upgrade - the first failure looked transient because the earlier run in the same hour passed on a cached step ordering. Threshold to >40 MB and a note that this is a truncation guard to resize on bumps, not a version pin. --- .github/workflows/ci.yml | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index f037dbe6..47a048dd 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -128,11 +128,13 @@ jobs: echo "code=${{ github.run_number }}" >> "$GITHUB_OUTPUT" - name: Fetch TTS runtime AAR - # sherpa-onnx neural-TTS + ASR runtime (Kokoro voice, Whisper voice search). It's a 57 MB + # sherpa-onnx neural-TTS + ASR runtime (Kokoro voice, Whisper voice search). It's a ~47 MB # prebuilt AAR (no Maven artifact) hosted on the fixed-tag `tts-runtime` release, gitignored # out of the repo - fetch it into app/libs/ so :app can package the arm64 AND armeabi-v7a # .so. Verify the size so a truncated download fails loud instead of producing a - # runtime-crashing APK. + # runtime-crashing APK. (The threshold is sized for 1.13.4's 48,847,529 bytes; the 1.13.3 + # AAR was 57 MB, and the old >50 MB check failed the UPGRADED artifact - resize this when + # the AAR is bumped, it is a truncation guard, not a version pin.) run: | mkdir -p app/libs # Hosted on this repo's fixed-tag `tts-runtime` release (a one-time manual upload). @@ -140,7 +142,7 @@ jobs: # armv7 unaligned-read SIGBUS that crashed every model load on 32-bit phones (issue #95). curl -fSL -o app/libs/sherpa-onnx-1.13.4.aar \ "https://github.com/${{ github.repository }}/releases/download/tts-runtime/sherpa-onnx-1.13.4.aar" - test "$(stat -c%s app/libs/sherpa-onnx-1.13.4.aar)" -gt 50000000 + test "$(stat -c%s app/libs/sherpa-onnx-1.13.4.aar)" -gt 40000000 # The hosted AAR must actually CARRY 32-bit ARM, not just be big enough. If it is ever # replaced with an arm64-only build, the v7a strip fix (#81) silently reverts: the APK # still builds, still installs on a TCL Flip 2, and voice search + Vela voice are dead