An indigenous physical AI robotics agent powered by Raspberry Pi 5, Google Gemini Live API, Faster-Whisper edge wake engine, and ESP32 dual-OLED expressive eyes.
WALL-E is an embodied AI robotics platform engineered to bring real-time multimodal intelligence into the physical world. Unlike pure software chatbots, WALL-E combines edge audio processing, offline wake word detection, dynamic Voice Activity Detection (VAD), physical facial expressions on dual OLED displays, and serial motor navigation with the reasoning power of Google's Gemini Live API.
┌──────────────────────────────────────────────┐
│ Raspberry Pi 5 (8GB) │
│ │
[ USB Microphone ] ─────────►│ AudioPipeline │
│ ├── Virtual Split Queue │
│ └── Dynamic Noise Floor VAD │
│ │
│ WakePipeline (Faster-Whisper) │
│ └── Wake Word: "Hey WALL-E" / "WALL-E" │
│ │
│ IdentityPipeline │
│ ├── ECAPA-TDNN Voice Biometrics │
│ └── OpenCV / dlib Face Engine │
│ │
│ WalleSession (Gemini Live Full-Duplex) │
│ ├── WebSockets Bidirectional Audio Stream │
│ ├── Context Memory & Behavioral State │
│ └── Live Tool Invocation │
│ │
[ USB / 3.5mm Speaker ] ◄────┤ ALSA / PipeWire Software Debounced Audio │
└──────┬───────────────────────────────┬───────┘
│ UART Serial │ USB Serial
▼ ▼
┌────────────────────────────┐ ┌────────────────────────────┐
│ ESP32 Emotion Engine │ │ Arduino Motor Controller │
│ │ │ │
│ Dual 128x64 I2C OLED Eyes │ │ Differential Drive Motors │
│ (Blink, Happy, Curious, │ │ (Forward, Turn, Spin, │
│ Sleep, Surprise, Neutral) │ │ Obstacle Avoidance) │
└────────────────────────────┘ └────────────────────────────┘
- Edge Wake Word Pipeline: Low-latency offline wake word detection using
faster-whisperwith dynamic noise floor calibration to prevent audio hallucinations on USB microphones. - Full-Duplex Gemini Live Streaming: Bidirectional raw PCM audio streaming over low-latency WebSockets with sub-second response times.
- Hardware Noise Debounce & AEC: Custom software-level Acoustic Echo Cancellation (AEC) debouncing and ALSA audio loop suppression, eliminating audio feedback loops without hardware AEC chips.
- Dual OLED Emotion Eyes: ESP32 microcontroller controlling dual 0.96" SSD1306 OLED displays rendering organic procedural eye animations (blinking, squinting, looking around, expressions) in sync with AI agent state transitions.
- Persistent Behavioral Memory: Lightweight SQLite engine retaining user profiles, cross-session conversational memory, and emotional context while staying within strict token budgets.
- Differential Drive Navigation: Serial interface to Arduino motor drivers supporting offline voice navigation commands (via Vosk speech recognition) and autonomous obstacle avoidance.
- Systemd Production Daemon: Auto-starting Linux daemon (
walle.service) with watchdog health monitoring and graceful USB hotplug recovery.
| ESP32 Pin | OLED 1 (Left Eye) | OLED 2 (Right Eye) | Raspberry Pi 5 |
|---|---|---|---|
| 3.3V | VCC | VCC | — |
| GND | GND | GND | GND (Pin 6) |
| GPIO 21 (SDA) | SDA | SDA (0x3D) | — |
| GPIO 22 (SCL) | SCL | SCL | — |
| GPIO 16 (RX2) | — | — | TXD / GPIO 14 (Pin 8) |
| GPIO 17 (TX2) | — | — | RXD / GPIO 15 (Pin 10) |
- Host OS: Raspberry Pi OS Bookworm 64-bit
- Core Orchestrator: Python 3.11,
asyncio,sounddevice,numpy,scipy - AI / LLM: Google Gemini Live API (WebSocket bidirectional audio)
- Local Edge ML: Faster-Whisper, Vosk, OpenCV
- Microcontrollers: ESP32 (C++/Arduino), Arduino Uno/Nano
- Audio Routing: ALSA, PipeWire, WebRTC AEC
- Database: SQLite3
git clone https://github.com/SahilKumar337/WALLE-AI-Robot.git ~/walle
cd ~/walle
python3 -m venv venv --system-site-packages
source venv/bin/activate
pip install -r requirements-pi.txtcp .env.example .env
nano .env # Add GEMINI_API_KEYpython3 main.pysudo cp deploy/pi/walle.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now walle.service├── main.py # Master hardware orchestrator
├── server.py # Optional FastAPI web visualizer
├── core/
│ ├── config.py # Central settings & hardware pins
│ ├── logger.py # Multi-level structured logger
│ ├── registry.py # Async pipeline lifecycle manager
│ └── tts.py # Local fallback TTS & audio chimes
├── pipelines/
│ ├── audio_pipeline.py # Dual-channel mic capture & dynamic VAD
│ ├── wake_pipeline.py # Faster-whisper offline wake detection
│ ├── gemini_pipeline.py # Google Gemini Live WebSocket session
│ ├── identity_pipeline.py # Voice biometrics & face recognition
│ ├── navigation_pipeline.py # Vosk offline commands & serial routing
│ └── tool_pipeline.py # Function calling & external tools
├── arduino_firmware/
│ ├── walle_eyes/ # ESP32 dual OLED animation firmware
│ └── robot_control.ino # Arduino motor control sketch
├── deploy/
│ └── pi/ # Systemd service & deployment scripts
└── scripts/ # Audio setup & ALSA/PipeWire configuration
Sahil Kumar
- GitHub: @SahilKumar337
- LinkedIn: linkedin.com/in/sahilkumar337
- Portfolio: sahil-kumar.vercel.app