English | 日本語
Vokra is a speech-first inference runtime written in Rust. It implements the audio pieces that general-purpose graph runtimes often leave to application code: streaming state, STFT/iSTFT and mel frontends, vocoders, neural codecs, CTC/RNN-T decoding, VAD, speaker features, pitch extraction, and audio enhancement.
Vokra loads provenance-aware GGUF files and does not load ONNX graphs at
runtime. The default runtime has no third-party Cargo dependencies: the root
Cargo.lock contains only first-party vokra-* crates.
Development status (2026-09-17): the workspace is
0.3.0development; the observedmainbaseline for this campaign is2b08b7f7. The latest live, read-only public inventory, evaluated with audit logic througha4af7800, reports 194 repositories, 193 GGUF-bearing repositories and 198 GGUF files: 136 Mac-CPU-complete, 43 partial, 14 without a runtime binder and 1 non-artifact; Apple Metal is 136 full, 57 blocked by CPU and 1 non-artifact. VibeVoice Realtime-0.5B now has a strict structuralvibevoice_streamingbinder and CLI inspection route, so it is partial; synthesis, the complete weight manifest, independent reference, and CPU parity remain pending. Qwen3-TTS PR #109 subsequently completed a clean, no-upload VAST run for all four public variants plus the shared 12 Hz decoder: strict official-weight reload and independent CPU parity passed 4/4, and the recovered batch closed 38/38 checksums. Apple CPU/Metal and no-fallback verification, and any public-artifact replacement, remain pending. The bounded Apple Silicon batch recorded in the Apple results passed only for its named scopes, and four explicitly approved artifacts were subsequently published. There are still 58 unresolved public rows. Version 0.3.0 is the first GitHub source release; external package registries remain disabled. Vokra remains pre-1.0, so Rust APIs, the C ABI, GGUF metadata and model coverage may change. Pin the exactv0.3.0tag when evaluating this release in another project.
- Audio-native execution: speech frontends, streaming caches, decoders, vocoders, codecs, VAD, and enhancement are native operators rather than ONNX graph glue.
- Small dependency surface: runtime crates depend only on first-party
vokra-*crates. Offline conversion remains separate from runtime loading. - Explicit failures: unsupported operations and unavailable devices return errors; GPU work never silently falls back to CPU.
- Reproducible model files: Vokra GGUF metadata records frontend settings, topology, quantization policy, source provenance, and licence information.
- Portable integration: CPU is the default; Metal, CUDA, Vulkan, and WebGPU are opt-in. A generated C header supports native and language bindings.
You need Git and Rust 1.89 or newer. Build the CLI from source:
git clone https://github.com/ayutaz/vokra.git
cd vokra
cargo build --release -p vokra-cliDownload the published Whisper base GGUF and run the included public-domain audio fixture:
curl -L https://huggingface.co/vokra/whisper-base/resolve/main/whisper-base.gguf \
-o whisper-base.gguf
target/release/vokra-cli run \
--model whisper-base.gguf \
--input tests/fixtures/audio/jfk-30s.wavUse the built-in help before converting or running another architecture:
target/release/vokra-cli --help
target/release/vokra-cli convert --help
target/release/vokra-cli run --helpThe getting-started guide covers conversion, VAD, TTS, benchmarking, and the C ABI.
Vokra covers ASR, TTS, speech-to-speech, VAD and turn-taking, keyword spotting, speaker processing, pitch, codecs and vocoders, enhancement, separation, and audio understanding. Maturity is tracked per architecture: a converter, a GGUF loader, a native forward pass, numerical parity, and a published artifact are separate milestones. The existence of one does not imply the others.
Use these sources instead of a copied model list:
vokra-cli convert --help— accepted converter identifiers;vokra-cli run --help— CLI-routed inputs, outputs, and backend options;- the Vokra model hub — published artifacts and model-specific licence cards;
crates/vokra-cli/src/engine.rs— explicit runtime routing and deferred-operation registry for developers.
CPU is the default backend. Metal, CUDA, Vulkan, and WebGPU are opt-in and have operation-specific coverage. CoreML has an experimental whole-submodel delegate path for the Whisper encoder; QNN remains an SDK-gated experimental delegate scaffold. See the backend guide before selecting an accelerator.
The current coverage snapshot separates source/runtime readiness from numerical
evidence and publication. Incomplete routes fail closed; unsupported operations
never fall back silently to CPU. The UTMOS legacy Lightning checkpoint is also
intentionally refused by the restricted weights_only=True loader, so no
UTMOS numeric parity result is claimed.
Build the C library with:
cargo build --release -p vokra-capiinclude/vokra.h is the generated C reference. The
API index links to the Rust and binding surfaces,
including Python, Swift/iOS, Unity, Godot, Android, web, and server examples.
The C ABI remains pre-1.0 and is not frozen.
- Documentation map
- Getting started
- CLI tutorial
- Architecture
- Backend guide
- Migration guide
- Licence audit and legal/compliance notes
Contributions are welcome. Read CONTRIBUTING.md before a large change, and use the good first tasks for scoped entry points. Bugs and proposals can be filed in GitHub Issues. Participation is governed by the Code of Conduct. Report vulnerabilities privately as described in the Security Policy, not in a public issue.
Vokra source code is licensed under Apache-2.0. Model weights and
reference assets may use different licences; review each model card,
docs/license-audit.md, and NOTICE before
redistribution or commercial use. Non-commercial weights are excluded from
the default publication path unless an explicit research-only gate is used.