Skip to content

Latest commit

Β 

History

41 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸ€– Gemma Offline AI - KMP

Kotlin Compose Multiplatform Platforms License

A Kotlin Multiplatform (KMP) application that runs Google's Gemma LLM completely offline on Android and iOS devices. Built with Compose Multiplatform for a shared UI experience, powered by LiteRT-LM on both platforms for unified on-device AI inference.

Your conversations stay private. No internet required. 100% on-device AI.


✨ Features

Feature Description
πŸ”’ Fully Offline Run Gemma models completely on-device without internet connection
πŸ“± Cross-Platform Single codebase for Android & iOS using Compose Multiplatform
πŸ’¬ Real-time Streaming See AI responses as they're generated token by token
πŸ“Ž Attachments Attach images and PDFs to your messages
🎨 Modern UI Beautiful Material 3 design with dark/light theme support
βš™οΈ Configurable Adjust temperature, max tokens, and top-p parameters
πŸ’Ύ Model Management Import, load, and manage multiple models
πŸš€ Native Performance Platform-specific optimizations via LiteRT-LM

πŸ“Έ Screenshots

Chat Screen Settings Dark Mode
Chat Interface Model Management Theme Support

πŸ—οΈ Architecture

The app follows a clean architecture pattern with expect/actual mechanism for platform-specific implementations:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                        Compose Multiplatform UI                  β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚  β”‚ ChatScreen   β”‚  β”‚ SettingsScreenβ”‚  β”‚ Components          β”‚  β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                         ViewModel Layer                          β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚ ChatViewModel (StateFlow, Coroutines)                    β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                         Domain Layer                             β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                 β”‚
β”‚  β”‚ ChatMessageβ”‚  β”‚ Attachment β”‚  β”‚ ModelConfigβ”‚                 β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                 β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                    Platform Abstraction (expect)                 β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                     β”‚
β”‚  β”‚ GemmaInference   β”‚  β”‚ AttachmentPicker β”‚                     β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                     β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚     Android (actual)   β”‚           iOS (actual)                  β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚  β”‚ LiteRT-LM SDK    β”‚  β”‚  β”‚ LiteRT-LM Swift SDK (SPM)        β”‚  β”‚
β”‚  β”‚ Engine           β”‚  β”‚  β”‚ Native Swift Bridge              β”‚  β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Project Structure

composeApp/src/
β”œβ”€β”€ commonMain/                    # Shared Kotlin code (95%+ shared)
β”‚   β”œβ”€β”€ domain/
β”‚   β”‚   β”œβ”€β”€ model/                 # ChatMessage, Attachment, ModelConfig, ModelState
β”‚   β”‚   └── repository/            # ModelRepository (expect)
β”‚   β”œβ”€β”€ inference/                 # GemmaInference (expect)
β”‚   β”œβ”€β”€ picker/                    # FilePicker, AttachmentPicker (expect)
β”‚   └── ui/
β”‚       β”œβ”€β”€ components/            # EmptyStateView, LoadingIndicator
β”‚       β”œβ”€β”€ screens/               # ChatScreen, SettingsScreen
β”‚       β”œβ”€β”€ theme/                 # Material 3 Theme, ExtendedColors
β”‚       └── viewmodel/             # ChatViewModel, ChatUiState
β”‚
β”œβ”€β”€ androidMain/                   # Android-specific implementations
β”‚   β”œβ”€β”€ inference/                 # GemmaInference.android.kt (LiteRT-LM)
β”‚   β”œβ”€β”€ picker/                    # FilePicker.android.kt, AttachmentPicker.android.kt
β”‚   └── repository/                # ModelRepository.android.kt
β”‚
└── iosMain/                       # iOS-specific implementations
    β”œβ”€β”€ inference/                 # GemmaInference.ios.kt (LiteRT-LM)
    β”œβ”€β”€ picker/                    # FilePicker.ios.kt, AttachmentPicker.ios.kt
    └── repository/                # ModelRepository.ios.kt

βœ… Why LiteRT-LM over MediaPipe?

LiteRT-LM is now the inference layer for both Android and iOS, replacing MediaPipe completely across the project.

Unified SDK

  • Single inference stack across Android (Kotlin) and iOS (Swift)
  • No more platform divergence in the AI layer
  • Shared concepts, model handling, and streaming behavior across both apps

Performance improvements

  • Multi-Token Prediction (MTP) for speculative decoding and faster token generation
  • GPU acceleration with OpenCL on Android and Metal on iOS
  • NPU support on devices with neural processing units
  • Smarter caching for faster subsequent model load times

Modern API design

  • Android: First-class Kotlin coroutine Flow support for streaming with minimal bridging
  • iOS: Native Swift async/await and AsyncStream support
  • True push-based token streaming with no polling workarounds

Multi-modal support

  • Text, image, and audio inputs supported out of the box

Tool use / Function calling

  • Models can call defined Kotlin or Swift functions directly

Active development

  • MediaPipe LLM Inference API is superseded by LiteRT-LM
  • LiteRT-LM receives ongoing updates and platform improvements

Consistent model format

  • .litertlm models work across both Android and iOS

πŸš€ Getting Started

Prerequisites

Requirement Version
Android Studio Meerkat (2025.1.1) or later
Xcode 16.0+ (for iOS)
JDK 17+
Kotlin 2.4.0

1. Clone the Repository

git clone https://github.com/aspect-dev/OfflineAI-KMP.git
cd OfflineAI-KMP

2. Download a Gemma Model

Download a compatible model from Kaggle or Hugging Face.

Recommended: Use .litertlm models when available for the best cross-platform LiteRT-LM experience. .bin models are also supported.

Model Size Recommended For
gemma-2b-it-gpu-int4.bin ~1.4 GB Most devices
gemma-3n-E2B-it.litertlm ~1.8 GB Newer devices
gemma-7b-it-gpu-int4.bin ~4.5 GB High-end devices

3. Build & Run

Android

# Build debug APK
./gradlew :composeApp:assembleDebug

# Or run directly
./gradlew :composeApp:installDebug

iOS

# Install CocoaPods dependencies
cd iosApp
pod install
cd ..

# Build Kotlin framework
./gradlew :composeApp:linkPodDebugFrameworkIosArm64

Then open iosApp/iosApp.xcworkspace in Xcode and run.

4. Load a Model

  1. Launch the app
  2. Go to Settings (gear icon)
  3. Tap "Browse Files to Import Model"
  4. Select your downloaded .litertlm or .bin file
  5. The model will be copied to app storage and loaded

πŸ“± Platform Requirements

Platform Minimum Recommended Notes
Android API 24 (7.0) API 34+ 4GB+ RAM, GPU support preferred
iOS iOS 16.0 iOS 17+ iPhone 12+ / iPad Pro for best performance

Device Recommendations

  • Android: Pixel 7+, Samsung Galaxy S22+, or equivalent
  • iOS: iPhone 12 or newer, iPad Pro (M1/M2/M4)

βš™οΈ Configuration

Model Parameters

Parameter Range Default Description
Temperature 0.0 - 1.0 0.7 Controls randomness (lower = focused, higher = creative)
Max Tokens 256 - 4096 2048 Maximum response length
Top-p 0.0 - 1.0 0.9 Nucleus sampling threshold

Theme

The app automatically follows system theme preferences. Supports:

  • 🌞 Light Mode
  • πŸŒ™ Dark Mode

πŸ”§ Technical Details

Dependencies

Library Version Platform Purpose
Compose Multiplatform 1.11.0 Both Shared UI framework
LiteRT-LM 0.11.0 Android On-device LLM inference
LiteRT-LM Swift 0.13.1+ iOS On-device LLM inference (GPU/Metal, streaming)
Kotlinx Coroutines 1.11.0 Both Async operations & Flow
Kotlinx Serialization 1.11.0 Both JSON serialization
Lifecycle ViewModel 2.10.0 Both MVVM architecture
Navigation Compose 2.10.0 Both Screen navigation

iOS Setup

LiteRT-LM is declared directly in iosApp.xcodeproj/project.pbxproj via Swift Package Manager β€” no manual Xcode steps required. When you open iosApp.xcworkspace, Xcode resolves and downloads LiteRT-LM automatically.

Package: https://github.com/google-ai-edge/LiteRT-LM
Version:  from 0.13.1 (upToNextMajorVersion)
Product:  LiteRTLM

CocoaPods is still used only for integrating the composeApp Kotlin framework. The Podfile no longer includes MediaPipe dependencies:

source 'https://cdn.cocoapods.org'

platform :ios, '16.0'
use_frameworks! :linkage => :static

target 'iosApp' do
  pod 'composeApp', :path => '../composeApp'
end

🀝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/AmazingFeature)
  3. Commit your changes (git commit -m 'Add some AmazingFeature')
  4. Push to the branch (git push origin feature/AmazingFeature)
  5. Open a Pull Request

πŸ“„ License

This project is open source under the MIT License. See LICENSE for details.

Note: Gemma models are subject to Google's Gemma Terms of Use.


πŸ™ Acknowledgments


Made with ❀️ using Kotlin Multiplatform

Learn more about Kotlin Multiplatform

About

No description or website provided.

Topics

Resources

Stars

66 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages