A Kotlin Multiplatform (KMP) application that runs Google's Gemma LLM completely offline on Android and iOS devices. Built with Compose Multiplatform for a shared UI experience, powered by LiteRT-LM on both platforms for unified on-device AI inference.
Your conversations stay private. No internet required. 100% on-device AI.
| Feature | Description |
|---|---|
| π Fully Offline | Run Gemma models completely on-device without internet connection |
| π± Cross-Platform | Single codebase for Android & iOS using Compose Multiplatform |
| π¬ Real-time Streaming | See AI responses as they're generated token by token |
| π Attachments | Attach images and PDFs to your messages |
| π¨ Modern UI | Beautiful Material 3 design with dark/light theme support |
| βοΈ Configurable | Adjust temperature, max tokens, and top-p parameters |
| πΎ Model Management | Import, load, and manage multiple models |
| π Native Performance | Platform-specific optimizations via LiteRT-LM |
| Chat Screen | Settings | Dark Mode |
|---|---|---|
| Chat Interface | Model Management | Theme Support |
The app follows a clean architecture pattern with expect/actual mechanism for platform-specific implementations:
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Compose Multiplatform UI β
β ββββββββββββββββ ββββββββββββββββ ββββββββββββββββββββββββ β
β β ChatScreen β β SettingsScreenβ β Components β β
β ββββββββββββββββ ββββββββββββββββ ββββββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β ViewModel Layer β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β ChatViewModel (StateFlow, Coroutines) β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β Domain Layer β
β ββββββββββββββ ββββββββββββββ ββββββββββββββ β
β β ChatMessageβ β Attachment β β ModelConfigβ β
β ββββββββββββββ ββββββββββββββ ββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β Platform Abstraction (expect) β
β ββββββββββββββββββββ ββββββββββββββββββββ β
β β GemmaInference β β AttachmentPicker β β
β ββββββββββββββββββββ ββββββββββββββββββββ β
ββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββββββ€
β Android (actual) β iOS (actual) β
β ββββββββββββββββββββ β ββββββββββββββββββββββββββββββββββββ β
β β LiteRT-LM SDK β β β LiteRT-LM Swift SDK (SPM) β β
β β Engine β β β Native Swift Bridge β β
β ββββββββββββββββββββ β ββββββββββββββββββββββββββββββββββββ β
ββββββββββββββββββββββββββ΄βββββββββββββββββββββββββββββββββββββββββ
composeApp/src/
βββ commonMain/ # Shared Kotlin code (95%+ shared)
β βββ domain/
β β βββ model/ # ChatMessage, Attachment, ModelConfig, ModelState
β β βββ repository/ # ModelRepository (expect)
β βββ inference/ # GemmaInference (expect)
β βββ picker/ # FilePicker, AttachmentPicker (expect)
β βββ ui/
β βββ components/ # EmptyStateView, LoadingIndicator
β βββ screens/ # ChatScreen, SettingsScreen
β βββ theme/ # Material 3 Theme, ExtendedColors
β βββ viewmodel/ # ChatViewModel, ChatUiState
β
βββ androidMain/ # Android-specific implementations
β βββ inference/ # GemmaInference.android.kt (LiteRT-LM)
β βββ picker/ # FilePicker.android.kt, AttachmentPicker.android.kt
β βββ repository/ # ModelRepository.android.kt
β
βββ iosMain/ # iOS-specific implementations
βββ inference/ # GemmaInference.ios.kt (LiteRT-LM)
βββ picker/ # FilePicker.ios.kt, AttachmentPicker.ios.kt
βββ repository/ # ModelRepository.ios.kt
LiteRT-LM is now the inference layer for both Android and iOS, replacing MediaPipe completely across the project.
- Single inference stack across Android (Kotlin) and iOS (Swift)
- No more platform divergence in the AI layer
- Shared concepts, model handling, and streaming behavior across both apps
- Multi-Token Prediction (MTP) for speculative decoding and faster token generation
- GPU acceleration with OpenCL on Android and Metal on iOS
- NPU support on devices with neural processing units
- Smarter caching for faster subsequent model load times
- Android: First-class Kotlin coroutine
Flowsupport for streaming with minimal bridging - iOS: Native Swift
async/awaitandAsyncStreamsupport - True push-based token streaming with no polling workarounds
- Text, image, and audio inputs supported out of the box
- Models can call defined Kotlin or Swift functions directly
- MediaPipe LLM Inference API is superseded by LiteRT-LM
- LiteRT-LM receives ongoing updates and platform improvements
.litertlmmodels work across both Android and iOS
| Requirement | Version |
|---|---|
| Android Studio | Meerkat (2025.1.1) or later |
| Xcode | 16.0+ (for iOS) |
| JDK | 17+ |
| Kotlin | 2.4.0 |
git clone https://github.com/aspect-dev/OfflineAI-KMP.git
cd OfflineAI-KMPDownload a compatible model from Kaggle or Hugging Face.
Recommended: Use
.litertlmmodels when available for the best cross-platform LiteRT-LM experience..binmodels are also supported.
| Model | Size | Recommended For |
|---|---|---|
gemma-2b-it-gpu-int4.bin |
~1.4 GB | Most devices |
gemma-3n-E2B-it.litertlm |
~1.8 GB | Newer devices |
gemma-7b-it-gpu-int4.bin |
~4.5 GB | High-end devices |
# Build debug APK
./gradlew :composeApp:assembleDebug
# Or run directly
./gradlew :composeApp:installDebug# Install CocoaPods dependencies
cd iosApp
pod install
cd ..
# Build Kotlin framework
./gradlew :composeApp:linkPodDebugFrameworkIosArm64Then open iosApp/iosApp.xcworkspace in Xcode and run.
- Launch the app
- Go to Settings (gear icon)
- Tap "Browse Files to Import Model"
- Select your downloaded
.litertlmor.binfile - The model will be copied to app storage and loaded
| Platform | Minimum | Recommended | Notes |
|---|---|---|---|
| Android | API 24 (7.0) | API 34+ | 4GB+ RAM, GPU support preferred |
| iOS | iOS 16.0 | iOS 17+ | iPhone 12+ / iPad Pro for best performance |
- Android: Pixel 7+, Samsung Galaxy S22+, or equivalent
- iOS: iPhone 12 or newer, iPad Pro (M1/M2/M4)
| Parameter | Range | Default | Description |
|---|---|---|---|
| Temperature | 0.0 - 1.0 | 0.7 | Controls randomness (lower = focused, higher = creative) |
| Max Tokens | 256 - 4096 | 2048 | Maximum response length |
| Top-p | 0.0 - 1.0 | 0.9 | Nucleus sampling threshold |
The app automatically follows system theme preferences. Supports:
- π Light Mode
- π Dark Mode
| Library | Version | Platform | Purpose |
|---|---|---|---|
| Compose Multiplatform | 1.11.0 | Both | Shared UI framework |
| LiteRT-LM | 0.11.0 | Android | On-device LLM inference |
| LiteRT-LM Swift | 0.13.1+ | iOS | On-device LLM inference (GPU/Metal, streaming) |
| Kotlinx Coroutines | 1.11.0 | Both | Async operations & Flow |
| Kotlinx Serialization | 1.11.0 | Both | JSON serialization |
| Lifecycle ViewModel | 2.10.0 | Both | MVVM architecture |
| Navigation Compose | 2.10.0 | Both | Screen navigation |
LiteRT-LM is declared directly in iosApp.xcodeproj/project.pbxproj via Swift Package Manager β no manual Xcode steps required. When you open iosApp.xcworkspace, Xcode resolves and downloads LiteRT-LM automatically.
Package: https://github.com/google-ai-edge/LiteRT-LM
Version: from 0.13.1 (upToNextMajorVersion)
Product: LiteRTLM
CocoaPods is still used only for integrating the composeApp Kotlin framework. The Podfile no longer includes MediaPipe dependencies:
source 'https://cdn.cocoapods.org'
platform :ios, '16.0'
use_frameworks! :linkage => :static
target 'iosApp' do
pod 'composeApp', :path => '../composeApp'
endContributions are welcome! Please feel free to submit a Pull Request.
- Fork the repository
- Create your feature branch (
git checkout -b feature/AmazingFeature) - Commit your changes (
git commit -m 'Add some AmazingFeature') - Push to the branch (
git push origin feature/AmazingFeature) - Open a Pull Request
This project is open source under the MIT License. See LICENSE for details.
Note: Gemma models are subject to Google's Gemma Terms of Use.
- Google Gemma - The on-device LLM
- LiteRT-LM - On-device ML inference framework (Android & iOS)
- Kotlin Multiplatform - Cross-platform development
- Compose Multiplatform - Shared UI framework
- Google AI Edge Gallery - UI inspiration
Made with β€οΈ using Kotlin Multiplatform