From 50e3b7abb3f94c0028a4cfcb76c527c4439847b3 Mon Sep 17 00:00:00 2001 From: alltechdev Date: Mon, 20 Jul 2026 21:06:56 -0400 Subject: [PATCH] Use a lot less memory, and give it back when the system asks Reported on a TCL Flip 2 as "the whole app is a bit slow" (#83). Measured on an M5 (2.9 GB, Android 13, standardDebug, median of 5 cold starts) the app held 831 MB at peak, sat at 421 MB idle, and gave back nothing at all when the OS asked it to shrink, because nothing in the tree implemented memory-pressure handling. The biggest win needs no device gate. The on-device speech model costs ~267 MB while loaded (~101 MB of weights plus ~146 MB of onnxruntime arena) and was kept for the whole process on the chance of a mic tap many users never make. It is now dropped after two minutes unused and rebuilt on next use, so the instant first tap that was asked for on 2026-07-10 is kept while the session-long hold is not. Device-verified: scudo:secondary 111 MB to 9 MB at the 120 s mark, model rebuilding correctly afterwards. Idle PSS 421 MB to 299 MB on a phone that is not low-RAM at all. Nothing released under pressure. There was no onTrimMemory, onLowMemory or ComponentCallbacks2 anywhere, so a TRIM_MEMORY_COMPLETE freed 0 KB and the system's only remaining option was to kill us. VelaApp now fans every trim out through a new MemoryPressure holder, and the things that actually hold memory register a release: the speech model, the neural voice, MapLibre's native tile and sprite caches (MapView.onLowMemory was never called), all five hidden WebViews, and the image cache. WhisperRecognizer had no release path at all, so even Remove-model left ~267 MB resident for the rest of the process. Nothing adapted to the device either. There was no isLowRamDevice branch and the image cache was a flat 48 MB whatever the phone. Constrained devices now skip the startup preload of the speech model, cap images at 16 MB, skip the speculative WebView warm on every search, and fetch 8 ambient POI category terms instead of 15 with a smaller result pool. Roomier phones keep their existing behaviour. The low-RAM POI subset deliberately keeps school and park: the ambient layer filter-hides the basemap OSM poi layers at z14+, so those two have no second source and a first 6-term subset made every park and school pin vanish. Caught by an A/B screenshot, not by any test. Also stop shipping x86 and x86_64 native libraries. No phone Vela targets can execute them and libmaplibre.so alone carried 23 MB of them into every install. armeabi-v7a stays, since 32-bit ARM keypad phones are real. Measured, main vs this, low-RAM path: peak 831 MB to 581 MB (-30%), post-trim 397 MB to 246 MB (-38%), native heap 223 MB to 95 MB (-57%), cold start 4811 ms to 4333 ms. On a normal-RAM device idle drops 29%, post-trim 28% and native heap 44%. APK 97.4 MB to 92.0 MB. Debug builds honour `setprop debug.vela.lowram true` so the low-RAM path can be exercised on a dev phone, where it is otherwise dead code (every device we own reports lowRam=false heapClassMb=256). AGENTS.md: document the seam and the measurement traps, and correct the memory rule, which said the OverpassTrafficSignals/OverpassPois stream-parse follow-up was pending. It has been done for some time; the remaining buffered hot reader is the Google ambient path, and chasing the stale line wasted a pass. --- AGENTS.md | 51 +++++++- app/build.gradle.kts | 12 +- app/src/main/java/app/vela/VelaApp.kt | 28 ++++- .../main/java/app/vela/ui/MemoryPressure.kt | 110 ++++++++++++++++++ .../main/java/app/vela/ui/map/MapViewModel.kt | 14 ++- .../main/java/app/vela/ui/map/VelaMapView.kt | 10 ++ .../main/java/app/vela/voice/PiperSynth.kt | 13 +++ .../java/app/vela/voice/WhisperRecognizer.kt | 103 +++++++++++++++- .../java/app/vela/web/WebDirectionsFetcher.kt | 20 +++- .../main/java/app/vela/web/WebPhotoFetcher.kt | 19 +++ .../app/vela/web/WebPopularTimesFetcher.kt | 22 +++- .../java/app/vela/web/WebReviewsFetcher.kt | 20 +++- .../app/vela/web/WebStopDeparturesFetcher.kt | 20 +++- .../java/app/vela/core/data/LowRamMode.kt | 17 +++ .../core/data/google/GoogleMapsDataSource.kt | 29 ++++- docs/FEATURES.md | 2 + 16 files changed, 458 insertions(+), 32 deletions(-) create mode 100644 app/src/main/java/app/vela/ui/MemoryPressure.kt create mode 100644 core/src/main/java/app/vela/core/data/LowRamMode.kt diff --git a/AGENTS.md b/AGENTS.md index 7ea33030..af526d06 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -627,9 +627,56 @@ state - upstream's own 13ac02e8 already made the layers panel a VelaMenu): (raises the ceiling ~2x); don't remove it. (2) **Any Overpass / large-HTTP-body reader MUST stream-parse** - `Json.decodeFromStream(body.byteStream())` into a tiny `@Serializable` DTO, NEVER `resp.body.string()` + `parseToJsonElement` (that held ~5-10x the wire size in transient heap and - OOM'd mid-read - the Flock `out body` fetch per pan did this; fixed in `OverpassAlprCameras`, - `OverpassTrafficSignals`/`OverpassPois` are the same pattern + a pending follow-up). And NEVER lower a + OOM'd mid-read - the Flock `out body` fetch per pan did this). And NEVER lower a per-viewport Overpass fetch's min-zoom without shrinking the box. + **The Overpass follow-up is DONE (audited 2026-07-20):** `OverpassAlprCameras`, `OverpassTrafficSignals` + AND `OverpassPois` all `decodeFromStream` today. This line previously said the latter two were pending, + which sent a low-RAM investigation chasing already-fixed code. The remaining fully-buffered hot reader + is the GOOGLE ambient path, not Overpass: `GoogleMapsDataSource.get()` ends in `.string()`, and + `GoogleResponse.parse` then makes a `substring` copy plus a full `JsonElement` DOM - times a 15-term + fan-out per pan. That, not Overpass, is the ~180 MB/12 s. It cannot simply `decodeFromStream`: the + payload is a positional nameless array walked by `at(0,1,3)` paths, so there is no DTO to decode into. + The levers that DO move it are the term count and the `!7i` pool size. + +- **Low-RAM devices are a FIRST-CLASS target, and the app now adapts to them (issue #83, 2026-07-20).** + D-pad-first means feature-phone-first, and those phones are memory-poor as well as small. + - **`app/ui/MemoryPressure.kt` is the one seam.** `init()` from `VelaApp` classifies the device + (`ActivityManager.isLowRamDevice` OR heap class <= 127 MB) and `VelaApp.onTrimMemory` fans every + `TRIM_MEMORY_*` out to registered holders. **Anything that allocates something large or NATIVE + must register a release callback.** Registration, never a Hilt entry point: reaching a singleton + from a trim would CONSTRUCT it, so the trim would allocate the very thing it is freeing. + - `LowRamMode.enabled` (`:core`) is the `:core`-visible mirror, pushed in by `VelaApp` - same seam + as `CategoryFilter.enabled`, because `:core` must never read an `:app` holder. + - **Verify the low-RAM path or it ships unverified.** Every dev phone we own reports + `lowRam=false heapClassMb=256`, so those branches are dead code locally. Debug builds honour + `adb shell setprop debug.vela.lowram true` (then relaunch); `setprop debug.vela.lowram false` + restores real detection. NB `setprop ""` is a syntax error, not a reset. + - **Measuring: `am send-trim-memory` REFUSES background levels on a foreground process** + ("Unable to set a background trim level on a foreground process"). Press HOME first. A harness + that discards that stderr measures NOTHING and reports a clean baseline - this happened here and + produced a whole benchmark of void numbers before the error was noticed. Always check it. + - Measured on an M5 (2.9 GB, Android 13, standardDebug, median of 5 cold starts) main vs the fix, + low-RAM path: peak PSS 831 MB -> 581 MB (-30%), post-trim 397 MB -> 246 MB (-38%), native heap + 223 MB -> 95 MB (-57%), cold start 4811 ms -> 4333 ms. Idle PSS run-to-run variance is +-60 MB, + so single idle readings prove nothing; compare post-trim, which is paired within a run. + - **The ASR model is the single largest reclaimable allocation: ~267 MB PSS** (~101 MB of weights + in `scudo:secondary` plus ~146 MB of onnxruntime arena in `scudo:primary`). It releases on a + severe trim, is NOT warmed at startup on a low-RAM device, and - on EVERY device - is dropped + after `REAP_IDLE_MS` (120 s) unused and rebuilt on next use. `WhisperRecognizer.release()` + declines while a listen is in flight - freeing the native recognizer under a running decode is a + use-after-free that takes the process down rather than throwing. + - **Do not reach for a device gate when an IDLE gate will do.** The warm-at-startup behaviour was + first made low-RAM-conditional, which protected the instant-first-mic-tap UX on roomier phones + but left them holding 267 MB all session. Reaping on idle keeps that UX AND reclaims the memory + everywhere: device-verified on the 2.9 GB M5, `scudo:secondary` 111 MB -> 9 MB at the 120 s mark + with the model rebuilding on next use. Idle PSS 421 MB -> 299 MB (-29%) on a NON-low-RAM device. + Ask "can this be released when unused?" before "which devices should get less?". + - **`MapView.onLowMemory()` must be called.** MapLibre's tile/glyph/sprite caches are native and + that is the only way to shrink them; nothing called it before. + - **When you gate a category fan-out down, check what has no SECOND source.** The low-RAM ambient + subset keeps `school` and `park` on purpose: the ambient layer filter-hides the basemap OSM poi + layers at z14+, so those two would vanish entirely. A first 6-term subset did exactly that and + was caught by an A/B screenshot, not by any test. ## Layout diff --git a/app/build.gradle.kts b/app/build.gradle.kts index a03ccff7..8194c01a 100644 --- a/app/build.gradle.kts +++ b/app/build.gradle.kts @@ -217,12 +217,18 @@ android { resources { excludes += "/META-INF/{AL2.0,LGPL2.1}" } // The neural-TTS runtime (ONNX Runtime + sherpa-onnx, from the vendored AAR) ships its .so // for all 4 ABIs; Vela targets arm64 phones, so drop the other ABIs' copies - they'd add - // ~65 MB for no device we support. MapLibre and other libs stay multi-ABI (untouched). + // ~65 MB for no device we support. + // + // x86/x86_64 go for EVERY native lib, not just the TTS ones (issue #83). Those two ABIs + // exist for emulators; no phone Vela targets can execute them, and libmaplibre.so alone + // shipped 11.6 MB of x86_64 plus 11.6 MB of x86 in every install. armeabi-v7a is KEPT: + // it is a real 32-bit ARM target and dropping it would silently strand those phones, + // which is the opposite of this issue's goal. jniLibs { excludes += listOf( "**/armeabi-v7a/libonnxruntime.so", "**/armeabi-v7a/libsherpa-onnx*.so", - "**/x86/libonnxruntime.so", "**/x86/libsherpa-onnx*.so", - "**/x86_64/libonnxruntime.so", "**/x86_64/libsherpa-onnx*.so", + "**/x86/*.so", + "**/x86_64/*.so", ) } } diff --git a/app/src/main/java/app/vela/VelaApp.kt b/app/src/main/java/app/vela/VelaApp.kt index 8b5fc595..db0a2c9c 100644 --- a/app/src/main/java/app/vela/VelaApp.kt +++ b/app/src/main/java/app/vela/VelaApp.kt @@ -39,15 +39,34 @@ class VelaApp : Application(), coil.ImageLoaderFactory { * ~128 MB of decoded gallery bitmaps by design, which is most of the "rapid place churn * runs into the ceiling" OOM (issue #182; measured: 3 gallery-bearing places grew the live * Dalvik heap 14 -> 94 MB). 48 MB still holds a couple of screens of thumbnails + a hero - * or two; everything else re-decodes from Coil's disk cache, which is untouched. */ + * or two; everything else re-decodes from Coil's disk cache, which is untouched. + * + * The cap is now a function of the device instead of one constant: a 48 MB bitmap cache is + * reasonable on a 2-3 GB phone and absurd on a keypad phone whose whole heap class is 96 MB + * (issue #83). Low-RAM devices get 16 MB, which still covers a screen of result thumbnails. + * [MemoryPressure.init] must run before this, and does - onCreate inits it first. */ override fun newImageLoader(): coil.ImageLoader = coil.ImageLoader.Builder(this) .memoryCache { coil.memory.MemoryCache.Builder(this) - .maxSizeBytes(48 * 1024 * 1024) + .maxSizeBytes(if (app.vela.ui.MemoryPressure.lowRam) 16 * 1024 * 1024 else 48 * 1024 * 1024) .build() } .build() + /** + * Hand OS memory pressure to every holder that owns a large or native allocation (issue #83). + * Before this existed nothing in the app implemented ComponentCallbacks2, so a + * TRIM_MEMORY_COMPLETE released nothing at all and the OS had no option but to kill us. + * Coil's own cache is trimmed here; everything else releases through [MemoryPressure]. + */ + override fun onTrimMemory(level: Int) { + super.onTrimMemory(level) + app.vela.ui.MemoryPressure.dispatch(level) + if (app.vela.ui.MemoryPressure.isSevere(level)) { + runCatching { coil.Coil.imageLoader(this).memoryCache?.clear() } + } + } + /** Apply the persisted in-app language to the Application context too (no-op when following the * system), so `getString` from the ViewModel/nav-notification also localizes - resolved at launch * from the saved pref (an in-session change re-reads it on next launch). */ @@ -64,6 +83,11 @@ class VelaApp : Application(), coil.ImageLoaderFactory { Timber.plant(DiagTree(diag)) if (BuildConfig.DEBUG) Timber.plant(Timber.DebugTree()) + // Device memory class first: the Coil cap and the eager-warm decisions below both read it. + app.vela.ui.MemoryPressure.init(this) + // Push the device class down to :core, which cannot read an :app holder (same seam as + // CategoryFilter.enabled). Gates the ambient POI fan-out in GoogleMapsDataSource. + app.vela.core.data.LowRamMode.enabled = app.vela.ui.MemoryPressure.lowRam Units.init(this) AppTheme.init(this) AppLocale.init(this) // resolve the app language (system default) → drives the nav-text locale diff --git a/app/src/main/java/app/vela/ui/MemoryPressure.kt b/app/src/main/java/app/vela/ui/MemoryPressure.kt new file mode 100644 index 00000000..30e8a652 --- /dev/null +++ b/app/src/main/java/app/vela/ui/MemoryPressure.kt @@ -0,0 +1,110 @@ +package app.vela.ui + +import android.app.ActivityManager +import android.content.ComponentCallbacks2 +import android.content.Context +import java.util.concurrent.CopyOnWriteArrayList +import timber.log.Timber + +/** + * Process-wide memory-pressure fan-out, in the same shape as the other app-level holders + * (`TransitLayer`, `AppTheme`): `init()` from `VelaApp`, then anything holding a large or native + * allocation registers a release callback. + * + * Why registration and not a Hilt entry point: reaching `WhisperRecognizer`/`PiperSynth` from + * `onTrimMemory` through an EntryPoint would CONSTRUCT them if they had never been used, so a + * trim would allocate the very models it is trying to free. A holder registers only once it + * actually owns something worth releasing, so a trim can never create work. + * + * Measured on the M5 (2.9 GB, Android 13, standardDebug) before this existed: TRIM_MEMORY_COMPLETE + * released 0 KB, because nothing in the app implemented ComponentCallbacks2 at all. + */ +object MemoryPressure { + + /** A registered releaser. [level] is a `ComponentCallbacks2.TRIM_MEMORY_*` constant. */ + fun interface Listener { + fun release(level: Int) + } + + private val listeners = CopyOnWriteArrayList() + + /** + * True when the OS classes this device as low-RAM (`ActivityManager.isLowRamDevice`) OR its + * heap class is small enough that our normal budgets do not fit. The heap-class arm matters: + * plenty of cheap keypad phones do NOT set the low-RAM system property yet still hand out a + * 96 MB heap class, and those are exactly the phones this work is for. + */ + @Volatile var lowRam: Boolean = false + private set + + /** The device's normal (non-large) heap class in MB. 0 until [init]. */ + @Volatile var heapClassMb: Int = 0 + private set + + fun init(context: Context) { + val am = context.getSystemService(Context.ACTIVITY_SERVICE) as? ActivityManager + heapClassMb = am?.memoryClass ?: 0 + val forced = forcedLowRam() + lowRam = forced ?: ((am?.isLowRamDevice == true) || (heapClassMb in 1..127)) + Timber.i( + "MemoryPressure init lowRam=%b heapClassMb=%d forced=%s", + lowRam, heapClassMb, forced?.toString() ?: "no", + ) + } + + /** + * Debug-only override so the low-RAM path can be exercised on a normal dev phone: + * + * adb shell setprop debug.vela.lowram true # then relaunch the app + * adb shell setprop debug.vela.lowram "" # back to real detection + * + * Without this the low-RAM branches are dead code on every device we actually own (the M5 dev + * phone reports heapClassMb=256, lowRam=false), which means they would ship unverified. Returns + * null when unset or on a non-debug build, so release behaviour is untouched. + */ + private fun forcedLowRam(): Boolean? { + if (!app.vela.BuildConfig.DEBUG) return null + val v = runCatching { + @Suppress("PrivateApi") + val sp = Class.forName("android.os.SystemProperties") + sp.getMethod("get", String::class.java).invoke(null, "debug.vela.lowram") as? String + }.getOrNull() + return when (v?.lowercase()) { + "true", "1" -> true + "false", "0" -> false + else -> null + } + } + + /** Register [listener]; returns a handle whose `close()` unregisters. Safe to call any time. */ + fun register(listener: Listener): AutoCloseable { + listeners.add(listener) + return AutoCloseable { listeners.remove(listener) } + } + + /** + * Fan a trim out to every registered holder. Each listener is isolated: one throwing must not + * stop the rest from releasing, since under real pressure we want every byte we can get. + */ + fun dispatch(level: Int) { + Timber.i("MemoryPressure dispatch level=%d listeners=%d", level, listeners.size) + for (l in listeners) { + runCatching { l.release(level) } + .onFailure { Timber.w(it, "MemoryPressure listener failed") } + } + } + + /** + * The app is backgrounded or the OS is genuinely short of memory, so caches that only speed + * things up should go. Everything at or above this level is a "drop it" signal. + */ + fun isSevere(level: Int): Boolean = + level >= ComponentCallbacks2.TRIM_MEMORY_BACKGROUND || + level == ComponentCallbacks2.TRIM_MEMORY_RUNNING_CRITICAL || + level == ComponentCallbacks2.TRIM_MEMORY_RUNNING_LOW + + /** Only the harshest levels, where we drop things that cost real time to rebuild. */ + fun isCritical(level: Int): Boolean = + level >= ComponentCallbacks2.TRIM_MEMORY_COMPLETE || + level == ComponentCallbacks2.TRIM_MEMORY_RUNNING_CRITICAL +} diff --git a/app/src/main/java/app/vela/ui/map/MapViewModel.kt b/app/src/main/java/app/vela/ui/map/MapViewModel.kt index 004d4f4a..134eabf9 100644 --- a/app/src/main/java/app/vela/ui/map/MapViewModel.kt +++ b/app/src/main/java/app/vela/ui/map/MapViewModel.kt @@ -1180,8 +1180,14 @@ class MapViewModel @Inject constructor( // popular times AND the photo gallery land faster when the user taps a result // (both idempotent; the photo warm primes the renderer + HTTP/2 sockets + cache // so the first place page skips the cold start). - viewModelScope.launch { runCatching { webPopularTimes.prewarm() } } - runCatching { webPhotos.warm() } + // Skipped on low-RAM devices: each warm spins up a Chromium renderer SPECULATIVELY, on the + // guess that a search predicts a place tap. When memory is the scarce resource that trade is + // backwards - the user pays two renderers on every search whether or not they open anything + // (issue #83). Those phones build the WebView on first real use instead. + if (!app.vela.ui.MemoryPressure.lowRam) { + viewModelScope.launch { runCatching { webPopularTimes.prewarm() } } + runCatching { webPhotos.warm() } + } searchJob?.cancel() searchJob = viewModelScope.launch { // A fresh typed search leaves any along-route browse: picks open places normally again. @@ -4360,6 +4366,10 @@ class MapViewModel @Inject constructor( } fun deleteAsrModel() { + // Free the loaded model BEFORE removing its files. Deleting the directory alone left the + // native recognizer resident for the rest of the process (~267 MB measured, issue #83), so + // "Remove" reclaimed disk but no memory at all. + whisperRecognizer.release() app.vela.voice.AsrModel.dir(appContext).deleteRecursively() _state.update { it.copy(asrInstalled = false) } } diff --git a/app/src/main/java/app/vela/ui/map/VelaMapView.kt b/app/src/main/java/app/vela/ui/map/VelaMapView.kt index f7ed559d..e93dabe0 100644 --- a/app/src/main/java/app/vela/ui/map/VelaMapView.kt +++ b/app/src/main/java/app/vela/ui/map/VelaMapView.kt @@ -1084,7 +1084,17 @@ fun VelaMapView( } } lifecycleOwner.lifecycle.addObserver(observer) + // MapLibre keeps its tile, glyph and sprite caches in NATIVE memory, and the only way to ask + // it to shrink them is onLowMemory(). Nothing called it before (issue #83), so the map held + // its full cache through every trim the OS sent. Registered with the map's own lifecycle so + // the listener can never outlive the MapView it points at. + val trim = app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isSevere(level)) { + runCatching { mapView.onLowMemory() } + } + } onDispose { + trim.close() lifecycleOwner.lifecycle.removeObserver(observer) mapView.onPause() mapView.onStop() diff --git a/app/src/main/java/app/vela/voice/PiperSynth.kt b/app/src/main/java/app/vela/voice/PiperSynth.kt index 665bc37a..99f30f01 100644 --- a/app/src/main/java/app/vela/voice/PiperSynth.kt +++ b/app/src/main/java/app/vela/voice/PiperSynth.kt @@ -38,6 +38,19 @@ class PiperSynth @Inject constructor( @Volatile private var loadFailed = false @Volatile private var generation = 0 + init { + // The Piper VITS model is the app's second-largest native holding after the ASR model. + // CRITICAL only, deliberately narrower than the recognizer's severe trigger: dropping the + // synth costs a reload on the next prompt, and a prompt arriving late during navigation is + // a missed turn. TRIM_MEMORY_COMPLETE only reaches background processes, and + // RUNNING_CRITICAL means the device is about to start killing things regardless. + // release() posts to the piper-tts worker, so it is already serialized against an + // in-flight synthesis and cannot free the model out from under one (issue #83). + app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isCritical(level)) release() + } + } + /** Which voice id `tts` currently holds - lets [ensureLoaded] detect a voice switch and rebuild. */ @Volatile private var loadedVoiceId: String? = null diff --git a/app/src/main/java/app/vela/voice/WhisperRecognizer.kt b/app/src/main/java/app/vela/voice/WhisperRecognizer.kt index a7654d1e..1edf95f1 100644 --- a/app/src/main/java/app/vela/voice/WhisperRecognizer.kt +++ b/app/src/main/java/app/vela/voice/WhisperRecognizer.kt @@ -34,9 +34,11 @@ import kotlin.math.sqrt * the end of speech, and returns the transcript. Nothing leaves the phone and no third-party voice * app is needed (that's tier-2 - the RECOGNIZE_SPEECH intent handoff in MapScreen). * - * The Whisper recognizer loads lazily and is kept for the process lifetime (~1 s to load); the VAD is - * created per listen (it's tiny and holds streaming state). R8 must keep `com.k2fsa.sherpa.onnx.**` - * (JNI resolves classes by name) - already in `consumer-rules`/`proguard` for Piper. + * The Whisper recognizer loads lazily (~1 s) and is NO LONGER kept for the process lifetime: it costs + * ~267 MB PSS, so it is dropped after [REAP_IDLE_MS] of quiet and on any severe memory trim, and + * rebuilt on the next use (issue #83). The VAD is created per listen (it's tiny and holds streaming + * state). R8 must keep `com.k2fsa.sherpa.onnx.**` (JNI resolves classes by name) - already in + * `consumer-rules`/`proguard` for Piper. */ @Singleton class WhisperRecognizer @Inject constructor( @@ -46,6 +48,67 @@ class WhisperRecognizer @Inject constructor( @Volatile private var recognizer: OfflineRecognizer? = null @Volatile private var loadedLang: String? = null + /** Non-zero while a [listen] is inside the native recognizer. [release] refuses to free the + * model while this is set: `OfflineRecognizer.release()` frees C++ memory that an in-flight + * decode is still reading, which is a use-after-free that takes the process down rather than + * throwing. A trim arriving mid-utterance simply keeps the model until the utterance ends. */ + private val inFlight = java.util.concurrent.atomic.AtomicInteger(0) + + /** Idle-reap timer, same idea as the web fetchers' `REAP_IDLE_MS` (issue #182). One daemon + * thread, shared, created lazily so a device that never loads the model never starts it. */ + private val reaper by lazy { + java.util.concurrent.Executors.newSingleThreadScheduledExecutor { r -> + Thread(r, "asr-reaper").apply { isDaemon = true } + } + } + @Volatile private var reapTask: java.util.concurrent.ScheduledFuture<*>? = null + + init { + // Measured on an M5 (2.9 GB, standardDebug, issue #83): the loaded Whisper tiny int8 model + // costs ~267 MB PSS - ~101 MB of weights in scudo:secondary plus ~146 MB of onnxruntime + // arena in scudo:primary. That was resident for the whole process with no way to reclaim it, + // and it survived deleteAsrModel(). It is by far the largest single reclaimable allocation + // in the app, so it releases on any severe trim and reloads (~1 s) on the next listen. + app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isSevere(level)) release() + } + } + + /** + * Drop the model after a quiet period, on EVERY device, not just low-RAM ones. + * + * Warming at startup buys an instant first mic tap (a user asked for it, 2026-07-10) but the app + * was then holding ~267 MB for the whole session on the CHANCE of a tap that many users never + * make. Reaping after idle keeps the instant first tap and stops the model outliving the user's + * interest in it; a later tap pays the same ~1 s load the very first one used to. Every load and + * every listen re-arms the timer, so an active dictation session never reaps mid-use. + */ + private fun armIdleReap() { + reapTask?.cancel(false) + reapTask = runCatching { + reaper.schedule({ release() }, REAP_IDLE_MS, java.util.concurrent.TimeUnit.MILLISECONDS) + }.getOrNull() + } + + /** + * Free the native recognizer. Safe to call any time: no-op when nothing is loaded, and declines + * while a listen is in flight (see [inFlight]). The next [listen]/[warmUp] rebuilds it. + */ + fun release() { + if (inFlight.get() > 0) { + Timber.tag(TAG).i("release skipped, listen in flight") + return + } + synchronized(loadLock) { + val r = recognizer ?: return + recognizer = null + loadedLang = null + runCatching { r.release() } + .onFailure { Timber.tag(TAG).w(it, "recognizer release failed") } + Timber.tag(TAG).i("recognizer released") + } + } + private val audioManager by lazy { context.getSystemService(Context.AUDIO_SERVICE) as? AudioManager } @Volatile private var focusRequest: AudioFocusRequest? = null @@ -85,6 +148,10 @@ class WhisperRecognizer @Inject constructor( const val SAMPLE_RATE = 16000 const val VAD_WINDOW = 512 // Silero v4/v5 window at 16 kHz const val MAX_SECONDS = 15 // hard cap on one utterance + // Drop the loaded model after this quiet period (issue #83). Matches the web + // fetchers' REAP_IDLE_MS: long enough that a dictation session never reaps between + // utterances, short enough that a session-long 267 MB hold cannot happen. + const val REAP_IDLE_MS = 120_000L } fun isInstalled(): Boolean = @@ -123,6 +190,14 @@ class WhisperRecognizer @Inject constructor( * lazily on the next listen (rare enough not to chase). */ fun warmUp() { if (!AsrModel.isInstalled(context)) return + // On a low-RAM device the warm-up is a bad trade: it spends ~267 MB (measured, issue #83) at + // EVERY launch to save ~1 s on a mic tap the user may never make, and refreshAsr() calls this + // from VM init plus two LaunchedEffects. Those phones load on first listen instead. Roomier + // devices keep the instant-mic behaviour they have always had. + if (app.vela.ui.MemoryPressure.lowRam) { + Timber.tag(TAG).i("skipping ASR warm-up on a low-RAM device, will load on first listen") + return + } Thread({ runCatching { ensureRecognizer() } }, "asr-warmup").start() } @@ -184,6 +259,7 @@ class WhisperRecognizer @Inject constructor( prefs.edit().putBoolean(KEY_LOAD_INFLIGHT, false).apply() recognizer = r loadedLang = lang + if (r != null) armIdleReap() // start the quiet-period countdown from the load return r } } @@ -195,11 +271,30 @@ class WhisperRecognizer @Inject constructor( * (the user tapped done/close). Runs off the main thread; safe to cancel via coroutine too. */ /** Listen, transcribe, and say WHY when it does not work - see [VoiceResult]. Every failure exit - * logs under `VELAASR` so a tester's logcat names the cause without another round-trip. */ + * logs under `VELAASR` so a tester's logcat names the cause without another round-trip. + * + * Thin wrapper over [listenInner] that marks the recognizer busy, so a memory trim arriving + * mid-utterance cannot free the native model out from under the decode (see [inFlight]). The + * inner function keeps its many early returns; this keeps the guard exception-safe. */ suspend fun listen( onLevel: (Float) -> Unit, onListening: () -> Unit, cancelled: () -> Boolean, + ): VoiceResult { + inFlight.incrementAndGet() + reapTask?.cancel(false) // never reap mid-utterance + try { + return listenInner(onLevel, onListening, cancelled) + } finally { + inFlight.decrementAndGet() + armIdleReap() // restart the quiet period from the END of this utterance + } + } + + private suspend fun listenInner( + onLevel: (Float) -> Unit, + onListening: () -> Unit, + cancelled: () -> Boolean, ): VoiceResult = withContext(Dispatchers.Default) { fun fail(reason: VoiceResult.Reason, detail: String? = null): VoiceResult.Failed { Timber.tag(TAG).e("listen failed: $reason${detail?.let { " ($it)" } ?: ""}") diff --git a/app/src/main/java/app/vela/web/WebDirectionsFetcher.kt b/app/src/main/java/app/vela/web/WebDirectionsFetcher.kt index 98be24ff..fbf2b053 100644 --- a/app/src/main/java/app/vela/web/WebDirectionsFetcher.kt +++ b/app/src/main/java/app/vela/web/WebDirectionsFetcher.kt @@ -64,14 +64,26 @@ class WebDirectionsFetcher @Inject constructor( * real memory. The next fetch after a reap just re-creates it. */ private fun scheduleReap() { reap?.let(main::removeCallbacks) - val r = Runnable { - webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } - webView = null - } + val r = Runnable { reapNow() } reap = r main.postDelayed(r, REAP_IDLE_MS) } + /** Destroy the WebView immediately. Must run on the main thread (WebView requirement). */ + private fun reapNow() { + webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } + webView = null + } + + init { + // Under real memory pressure the 120 s idle timer is far too slow - the OS is asking for + // memory NOW and a Chromium renderer is one of the largest things we hold (issue #83). + // Reap on the main thread, since WebView.destroy() requires it. + app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { cancelReap(); reapNow() } + } + } + private fun cancelReap() { reap?.let(main::removeCallbacks) reap = null diff --git a/app/src/main/java/app/vela/web/WebPhotoFetcher.kt b/app/src/main/java/app/vela/web/WebPhotoFetcher.kt index 3e2b5669..7985f467 100644 --- a/app/src/main/java/app/vela/web/WebPhotoFetcher.kt +++ b/app/src/main/java/app/vela/web/WebPhotoFetcher.kt @@ -54,6 +54,25 @@ class WebPhotoFetcher @Inject constructor( @Volatile private var webView: WebView? = null @Volatile private var warmed = false + init { + // Unlike the other four web fetchers this one has NO idle reaper, so a warmed gallery + // renderer was pinned for the whole session with only renderer-death to clear it. A + // Chromium renderer is one of the largest things the app holds, and this fetcher is one of + // the two warmed speculatively on every search, so it releases under pressure (issue #83). + app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { reapNow() } + } + } + + /** Destroy the WebView immediately. Main thread only (WebView requirement). The next + * [warm]/fetch rebuilds it via `ensureWebView`, exactly as after a renderer death. */ + private fun reapNow() { + val wv = webView ?: return + webView = null + warmed = false + runCatching { wv.loadUrl("about:blank"); wv.destroy() } + } + // featureId → its scraped gallery. Re-tapping a place (or bouncing back from directions) then // shows photos INSTANTLY instead of re-running the ~20 s scrape. Access-order LRU, small cap. private val cache = object : LinkedHashMap>(16, 0.75f, true) { diff --git a/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt b/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt index dff125a7..90e77ae4 100644 --- a/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt +++ b/app/src/main/java/app/vela/web/WebPopularTimesFetcher.kt @@ -59,15 +59,27 @@ class WebPopularTimesFetcher @Inject constructor( * reap re-warms (google.com -> maps), a one-off few-second cost after minutes idle. */ private fun scheduleReap() { reap?.let(main::removeCallbacks) - val r = Runnable { - webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } - webView = null - warm = null // ensureWarm re-runs the warm sequence on the next fetch - } + val r = Runnable { reapNow() } reap = r main.postDelayed(r, REAP_IDLE_MS) } + /** Destroy the WebView immediately. Must run on the main thread (WebView requirement). */ + private fun reapNow() { + webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } + webView = null + warm = null // ensureWarm re-runs the warm sequence on the next fetch + } + + init { + // Under real memory pressure the 120 s idle timer is far too slow - the OS is asking for + // memory NOW and a Chromium renderer is one of the largest things we hold (issue #83). + // Reap on the main thread, since WebView.destroy() requires it. + app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { cancelReap(); reapNow() } + } + } + private fun cancelReap() { reap?.let(main::removeCallbacks) reap = null diff --git a/app/src/main/java/app/vela/web/WebReviewsFetcher.kt b/app/src/main/java/app/vela/web/WebReviewsFetcher.kt index d25bf7cc..051bb490 100644 --- a/app/src/main/java/app/vela/web/WebReviewsFetcher.kt +++ b/app/src/main/java/app/vela/web/WebReviewsFetcher.kt @@ -59,14 +59,26 @@ class WebReviewsFetcher @Inject constructor( * real memory. The next fetch after a reap just re-creates it. */ private fun scheduleReap() { reap?.let(main::removeCallbacks) - val r = Runnable { - webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } - webView = null - } + val r = Runnable { reapNow() } reap = r main.postDelayed(r, REAP_IDLE_MS) } + /** Destroy the WebView immediately. Must run on the main thread (WebView requirement). */ + private fun reapNow() { + webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } + webView = null + } + + init { + // Under real memory pressure the 120 s idle timer is far too slow - the OS is asking for + // memory NOW and a Chromium renderer is one of the largest things we hold (issue #83). + // Reap on the main thread, since WebView.destroy() requires it. + app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { cancelReap(); reapNow() } + } + } + private fun cancelReap() { reap?.let(main::removeCallbacks) reap = null diff --git a/app/src/main/java/app/vela/web/WebStopDeparturesFetcher.kt b/app/src/main/java/app/vela/web/WebStopDeparturesFetcher.kt index dc3a4edd..32d1dff4 100644 --- a/app/src/main/java/app/vela/web/WebStopDeparturesFetcher.kt +++ b/app/src/main/java/app/vela/web/WebStopDeparturesFetcher.kt @@ -54,14 +54,26 @@ class WebStopDeparturesFetcher @Inject constructor( * just re-creates it - a one-off warm-up, only after minutes of not using the feature. */ private fun scheduleReap() { reap?.let(main::removeCallbacks) - val r = Runnable { - webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } - webView = null - } + val r = Runnable { reapNow() } reap = r main.postDelayed(r, REAP_IDLE_MS) } + /** Destroy the WebView immediately. Must run on the main thread (WebView requirement). */ + private fun reapNow() { + webView?.let { runCatching { it.loadUrl("about:blank"); it.destroy() } } + webView = null + } + + init { + // Under real memory pressure the 120 s idle timer is far too slow - the OS is asking for + // memory NOW and a Chromium renderer is one of the largest things we hold (issue #83). + // Reap on the main thread, since WebView.destroy() requires it. + app.vela.ui.MemoryPressure.register { level -> + if (app.vela.ui.MemoryPressure.isSevere(level)) main.post { cancelReap(); reapNow() } + } + } + private fun cancelReap() { reap?.let(main::removeCallbacks) reap = null diff --git a/core/src/main/java/app/vela/core/data/LowRamMode.kt b/core/src/main/java/app/vela/core/data/LowRamMode.kt new file mode 100644 index 00000000..56d7272d --- /dev/null +++ b/core/src/main/java/app/vela/core/data/LowRamMode.kt @@ -0,0 +1,17 @@ +package app.vela.core.data + +/** + * Whether this device is memory-constrained, exposed as a `:core`-visible flag. + * + * Same shape and reason as [CategoryFilter.enabled]: the detection lives in `:app` + * (`app.vela.ui.MemoryPressure`, which needs `ActivityManager`), but the behaviour it gates has to + * act down at the data-source seam. `:core` stays UI-agnostic and never reads an app holder, so the + * app pushes the value in at startup instead. + * + * Off by default, which keeps every roomier device byte-identical to previous behaviour. + */ +object LowRamMode { + + /** Set once from `VelaApp.onCreate` after `MemoryPressure.init`. */ + @Volatile var enabled: Boolean = false +} diff --git a/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt b/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt index facacd4d..6f121bb9 100644 --- a/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt +++ b/core/src/main/java/app/vela/core/data/google/GoogleMapsDataSource.kt @@ -6,6 +6,7 @@ import app.vela.core.config.JsTransforms import app.vela.core.diag.DiagLog import app.vela.core.data.CalibrationNeededException import app.vela.core.data.CategoryFilter +import app.vela.core.data.LowRamMode import app.vela.core.data.MapDataSource import app.vela.core.data.RouteEngine import app.vela.core.data.RouteGeometry @@ -142,7 +143,7 @@ class GoogleMapsDataSource @Inject constructor( // FAN OUT across category terms + merge: one "places" query is biased to prominent food/ // shops, so it misses whole tiers (a strip mall's plumber, nail salon, IT shop). A handful // of category queries roughly DOUBLES local coverage (live: 22→52 unique within 600 m). - val terms = listOf( + val allTerms = listOf( "places", "restaurants", "coffee", "stores", "shopping", "services", "beauty salon", "fast food", // High-traffic everyday categories the food/shop-biased set above under-returns, so the map // shows a Google-like MIX (a gas station, a gym, a grocer) rather than mostly restaurants. @@ -154,12 +155,36 @@ class GoogleMapsDataSource @Inject constructor( // their prominence low, so they surface in quiet/residential views without crowding businesses. "school", "park", ) + // LOW-RAM: the fan-out is the app's single largest allocation burst. Each term buffers a + // full response String, a stripped copy, and a JsonElement DOM (GoogleResponse.parse), and + // 15 of those run 4-at-a-time per pan - the ~180 MB/12 s churn in AGENTS.md's memory rule. + // Constrained devices fetch an 8-term subset instead of 15. + // + // The subset is NOT just the first N. "school" and "park" are retained DELIBERATELY: while + // the ambient layer is active at z14+ the basemap's OSM poi layers are filter-hidden, so + // those categories have NO second source - dropping them makes parks and schools vanish + // from the map entirely, which is the exact bug the civic/green terms were added to fix + // (see the comment on allTerms above). Device-verified by A/B screenshot: a first attempt + // at a 6-term subset lost every park and school pin on the low-RAM frame. + // + // What goes instead are the terms whose places still surface via "places"/"stores" or whose + // absence degrades gracefully: shopping, services, beauty salon, fast food, gym, bar, + // pharmacy. Fewer ambient POIs is a visible trade, and the right one on a phone that + // otherwise OOMs (issue #83). Roomier devices are unaffected. + val terms = if (LowRamMode.enabled) { + listOf("places", "restaurants", "coffee", "stores", "grocery store", "gas station", "school", "park") + } else { + allTerms + } suspend fun fetchTerm(term: String): List = ambientFanout.withPermit { runCatching { val pb = SearchPb.build(term, center, cal.searchPb) .replaceFirst(Regex("!1d[0-9.]+"), "!1d${spanMeters.toInt()}") .replaceFirst(Regex("!4f[0-9.]+"), "!4f${String.format(java.util.Locale.US, "%.1f", zoom)}") - .replaceFirst(Regex("!7i\\d+"), "!7i60") // deep pool per term, so zooming in can go down the rank + // Deep pool per term, so zooming in can go down the rank. Halved on low-RAM: the + // pool size drives the RESPONSE BODY size, and the body is what gets buffered + // and DOM-parsed per term (issue #83). + .replaceFirst(Regex("!7i\\d+"), if (LowRamMode.enabled) "!7i30" else "!7i60") val url = "${cal.searchEndpoint}&q=${term.enc()}&pb=${pb.enc()}".localized() SearchParser.parse(term, GoogleResponse.parse(get(url)), center, cal.paths).places }.getOrDefault(emptyList()) diff --git a/docs/FEATURES.md b/docs/FEATURES.md index ec30bbf9..d4fe2235 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -284,6 +284,8 @@ Status legend: [x] done · [~] partial / in progress · [ ] planned - [x] CI builds, tests, signs and publishes a normal release `v0.0.` with debug and release APKs; no prerelease channel, tracked by Obtainium and the updater with zero config. - [x] **Opt-in diagnostics/debug export** (Settings → Diagnostics, off by default) - a local-only event log exportable to JSON via the share sheet, never auto-uploaded, wiped when turned off, in-memory only. - [x] **Crash/ANR/jank capture, all local** - an uncaught-exception handler persists stack traces + breadcrumbs to disk for export, ApplicationExitInfo harvests ANR/native/low-memory kills, a debug ANR watchdog and StrictMode flag stalls and main-thread I/O; captured even with diagnostics off, never auto-sent. +- [x] **Memory use cut across the board** (issue #83) - the on-device speech model costs ~267 MB while loaded and used to stay resident for the entire session; it is now dropped after two minutes unused and rebuilt on next use, which reclaims ~101 MB on ANY phone at no visible cost. Every large or native allocation also releases when the system reports memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and gave back nothing when asked. Measured on a 2.9 GB test device: idle memory down 29%, post-trim down 38%, native heap down 57%. +- [x] **Low-RAM device adaptation** (issue #83) - additionally, the app detects a memory-constrained phone (`isLowRamDevice` or a heap class of 128 MB or less) and adapts: the on-device speech model is not preloaded at startup, the image cache is capped at 16 MB instead of 48 MB, WebView renderers are not warmed speculatively, and the ambient POI fan-out drops from 15 category terms to 8 with a smaller result pool. Independently, every large or native allocation now releases on OS memory pressure - the speech model, the neural voice, MapLibre's native tile and sprite caches, all five hidden WebViews and the image cache - where previously the app implemented no memory-pressure handling at all and returned nothing when the OS asked. Measured on a 2.9 GB test device: peak memory down 30%, post-trim memory down 38%, native heap down 57%, cold start ~480 ms faster. x86 and x86_64 native libraries are no longer shipped (no target phone can run them), cutting 23 MB of dead weight from every install. - [x] **Trip recording + replay** (Settings → "Save my trips", off by default, separate opt-in) - records each drive's GPS trace to a local file replayable on the map at 3x through the real nav pipeline; saved on arrival, listed with Replay/Share/Delete, Share exporting the raw CSV; replay auto-routes to the destination. - [x] **Simulate driving (demo mode)** (Settings → Navigation, off by default) - Start drives any planned route as a synthetic GPS trace through the live-nav loop so nav runs anywhere for demos and screenshots; End stops it; turn off to navigate for real. - [x] **Simulate my location (demo mode)** (off by default) - Vela pretends you're at the map centre so the dot, directions origin and recenter read from there; turn off for real GPS.