AI Briefing
KO

voicebox: From Voice Cloning to Agent Conversations, All Locally Without the Cloud

jamiepine/voicebox

·2026.09.15 14:42

It unifies voice output and input features previously handled separately by ElevenLabs and WisprFlow. Clone a voice with just a few seconds of audio samples or choose from over 50 presets to generate natural-sounding speech in 23 languages. All processing is done locally, ensuring voice data is never sent to external servers.

Switch between 7 TTS engines, including Qwen3-TTS, Kokoro, and Chatterbox, depending on the situation. Selecting Chatterbox Turbo recognizes emotion tags like [laugh] and [sigh] to produce expressive dialogue. With 8 audio effects based on Spotify's pedalboard library and automatic chunking, even long scripts are synthesized seamlessly without interruptions.

Press a global hotkey to dictate text into any app, with Whisper-based STT converting it instantly. It includes a built-in MCP server, allowing AI agents like Claude Code or Cursor to respond using a specified voice. Built with Tauri and Rust, it delivers lightweight native performance compared to Electron.

GitHub
GitHub repository

jamiepine/voicebox

The open-source AI voice studio. Clone, dictate, create.

TypeScript

This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.