Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages.
-
Updated
May 4, 2026 - Python
Automatic video translator and dubber using Whisper, XTTS v2 for voice cloning, and Ollama for local LLM translation. Supports 100+ languages.
Professional local-first AI production pipeline for long-form narration. Clone voices and generate studio-grade audiobooks (M4B/MP3) using Coqui XTTS-v2 and support for Voxtral (cloud)
Book to MP3 converter. Convert e-books (FB2, EPUB, TXT) to MP3 audiobooks using various Text-to-Speech technologies.
High-performance Coqui TTS API server with a hybrid "Hot/Cold" worker architecture
Simple Gradio application integrated with Hugging Face Multimodals to support visual question answering chatbot and more features
Run XTTS with Docker/Podman for voice fine-tuning in Gradio's Web UI
Saya Voice Assistant for Discord AI voice bot: listens, detects keywords, chats via LM Studio, and replies with TTS or voice cloning.
VoxLibri: The Ultimate AI-Powered eBook to Audiobook Converter. 🎧📚 Transform any eBook into a high-quality audiobook with state-of-the-art neural Text-to-Speech (TTS) technology. Featuring voice cloning, multi-language support (English, Tamil, and more), and a sleek Streamlit UI for a premium narration experience.
Dubbing english videos into russian.
Production-grade neural Text-to-Speech desktop studio powered by Coqui XTTS v2 & PyTorch. Features long-form sentence segmentation, multi-speaker voice conditioning, and real-time bilingual (EN/AR) localization.
AI-powered multilingual Text-to-Speech web application built with FastAPI and XTTS v2. Supports English and Hindi speech generation, browser audio playback, and downloadable audio output.
FakeLess is an AI-driven audio deepfake prevention system that protects human voices from unauthorized AI voice cloning using adversarial machine learning techniques. Built with deep learning, speech processing, and cross-platform integration, it delivers proactive voice security through real-time defense training and robust anti-cloning pipelines.
19 speech engines + Whisper: Kokoro, XTTS v2, F5-TTS, Chatterbox, Fish Speech, Bark, Dia, Higgs Audio, Qwen Omni, VibeVoice, SpeechT5, Parler, OuteTTS, VITS, Edge, Voxtral, VoxCPM2, CSM and Orpheus. Voice cloning, SRT, audio editor, CLI/API. Offline Kokoro included. From portable-linux-in-a-box.
A Streamlit web app for AI-powered voice cloning using Coqui XTTS v2. Record or upload reference voices, clone speech in multiple languages, and generate natural audio outputs.
XTTS fine-tuning via CLI
A high-performance, multi-engine voice cloning studio and text-to-speech renderer supporting XTTS v2, Qwen3-TTS, Chatterbox TTS and RVC v2 with an isolated, conflict-free subprocess architecture.
Windows desktop app for document-to-speech, transcription, subtitles, media cleanup, and local voice modeling workflows.
This program is designed to provide a graphical user interface for the xtts_api_server project: https://github.com/daswer123/xtts-api-server
To associate your repository with the xtts-v2 topic, visit your repo's landing page and select "manage topics."