Skip to content

Latest commit

 

History

278 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Vokra

English | 日本語

CI License: Apache-2.0

Vokra is a speech-first inference runtime written in Rust. It implements the audio pieces that general-purpose graph runtimes often leave to application code: streaming state, STFT/iSTFT and mel frontends, vocoders, neural codecs, CTC/RNN-T decoding, VAD, speaker features, pitch extraction, and audio enhancement.

Vokra loads provenance-aware GGUF files and does not load ONNX graphs at runtime. The default runtime has no third-party Cargo dependencies: the root Cargo.lock contains only first-party vokra-* crates.

Development status (2026-09-17): the workspace is 0.3.0 development; the observed main baseline for this campaign is 2b08b7f7. The latest live, read-only public inventory, evaluated with audit logic through a4af7800, reports 194 repositories, 193 GGUF-bearing repositories and 198 GGUF files: 136 Mac-CPU-complete, 43 partial, 14 without a runtime binder and 1 non-artifact; Apple Metal is 136 full, 57 blocked by CPU and 1 non-artifact. VibeVoice Realtime-0.5B now has a strict structural vibevoice_streaming binder and CLI inspection route, so it is partial; synthesis, the complete weight manifest, independent reference, and CPU parity remain pending. Qwen3-TTS PR #109 subsequently completed a clean, no-upload VAST run for all four public variants plus the shared 12 Hz decoder: strict official-weight reload and independent CPU parity passed 4/4, and the recovered batch closed 38/38 checksums. Apple CPU/Metal and no-fallback verification, and any public-artifact replacement, remain pending. The bounded Apple Silicon batch recorded in the Apple results passed only for its named scopes, and four explicitly approved artifacts were subsequently published. There are still 58 unresolved public rows. Version 0.3.0 is the first GitHub source release; external package registries remain disabled. Vokra remains pre-1.0, so Rust APIs, the C ABI, GGUF metadata and model coverage may change. Pin the exact v0.3.0 tag when evaluating this release in another project.

Why Vokra

  • Audio-native execution: speech frontends, streaming caches, decoders, vocoders, codecs, VAD, and enhancement are native operators rather than ONNX graph glue.
  • Small dependency surface: runtime crates depend only on first-party vokra-* crates. Offline conversion remains separate from runtime loading.
  • Explicit failures: unsupported operations and unavailable devices return errors; GPU work never silently falls back to CPU.
  • Reproducible model files: Vokra GGUF metadata records frontend settings, topology, quantization policy, source provenance, and licence information.
  • Portable integration: CPU is the default; Metal, CUDA, Vulkan, and WebGPU are opt-in. A generated C header supports native and language bindings.

Quick start

You need Git and Rust 1.89 or newer. Build the CLI from source:

git clone https://github.com/ayutaz/vokra.git
cd vokra
cargo build --release -p vokra-cli

Download the published Whisper base GGUF and run the included public-domain audio fixture:

curl -L https://huggingface.co/vokra/whisper-base/resolve/main/whisper-base.gguf \
  -o whisper-base.gguf
target/release/vokra-cli run \
  --model whisper-base.gguf \
  --input tests/fixtures/audio/jfk-30s.wav

Use the built-in help before converting or running another architecture:

target/release/vokra-cli --help
target/release/vokra-cli convert --help
target/release/vokra-cli run --help

The getting-started guide covers conversion, VAD, TTS, benchmarking, and the C ABI.

Model and backend status

Vokra covers ASR, TTS, speech-to-speech, VAD and turn-taking, keyword spotting, speaker processing, pitch, codecs and vocoders, enhancement, separation, and audio understanding. Maturity is tracked per architecture: a converter, a GGUF loader, a native forward pass, numerical parity, and a published artifact are separate milestones. The existence of one does not imply the others.

Use these sources instead of a copied model list:

  • vokra-cli convert --help — accepted converter identifiers;
  • vokra-cli run --help — CLI-routed inputs, outputs, and backend options;
  • the Vokra model hub — published artifacts and model-specific licence cards;
  • crates/vokra-cli/src/engine.rs — explicit runtime routing and deferred-operation registry for developers.

CPU is the default backend. Metal, CUDA, Vulkan, and WebGPU are opt-in and have operation-specific coverage. CoreML has an experimental whole-submodel delegate path for the Whisper encoder; QNN remains an SDK-gated experimental delegate scaffold. See the backend guide before selecting an accelerator.

The current coverage snapshot separates source/runtime readiness from numerical evidence and publication. Incomplete routes fail closed; unsupported operations never fall back silently to CPU. The UTMOS legacy Lightning checkpoint is also intentionally refused by the restricted weights_only=True loader, so no UTMOS numeric parity result is claimed.

Library integration

Build the C library with:

cargo build --release -p vokra-capi

include/vokra.h is the generated C reference. The API index links to the Rust and binding surfaces, including Python, Swift/iOS, Unity, Godot, Android, web, and server examples. The C ABI remains pre-1.0 and is not frozen.

Documentation

Contributing

Contributions are welcome. Read CONTRIBUTING.md before a large change, and use the good first tasks for scoped entry points. Bugs and proposals can be filed in GitHub Issues. Participation is governed by the Code of Conduct. Report vulnerabilities privately as described in the Security Policy, not in a public issue.

Licence

Vokra source code is licensed under Apache-2.0. Model weights and reference assets may use different licences; review each model card, docs/license-audit.md, and NOTICE before redistribution or commercial use. Non-commercial weights are excluded from the default publication path unless an explicit research-only gate is used.

About

Speech-first inference runtime in Rust — TTS / ASR / speech-to-speech / VC / speaker ID / VAD. An ONNX Runtime alternative that loads GGUF & safetensors directly: zero external dependencies, C ABI for Unity & Godot, CPU / Metal / CUDA / Vulkan backends. Apache-2.0, pre-release.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

13 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages