The open-source voice studio Voicebox clones a voice locally from a short audio sample and lets MCP-compatible AI agents such as Claude Code and Cursor speak in that voice. It runs entirely on-device without the cloud, and as of v0.5.0 it has surpassed 1.5 million downloads.
Open Source · Voicebox v0.5.0
Your voice, cloned on-device — now your AI agents can speak in it
A free, fully local voice studio that replicates a voice from a few seconds of audio and pipes it to MCP-compatible agents like Claude Code and Cursor — no cloud, no monthly fees.
1.5M+
Downloads as of v0.5.0
3s
Sample needed to clone a voice
99
Languages for dictation (Whisper)
CPU synthesis speed — LuxTTS vs realtime
Same unit (× realtime). LuxTTS generates 150× faster than playback, on CPU alone.
150×
LuxTTS on CPU (48kHz, ~1GB VRAM)
How an AI agent gets your voice
1 · Sample
A few seconds of clear audio, cloned zero-shot on-device
→
2 · MCP server
Exposes a voicebox.speak tool to agents
→
3 · Agent speaks
Claude Code, Cursor & Cline reply in the cloned voice
Pick an engine by use case
Up to 23 languages · 50,000 chars per generation with auto split & crossfade
Qwen3-TTS (1.7B/0.6B)
High-quality multilingual with instruction control for tone, pace & emotion
Chatterbox Turbo
Paralinguistic tags like [laugh], [sigh] — and fast
LuxTTS
150× realtime on CPU, 48kHz, ~1GB VRAM
Kokoro (82M)
Realtime on CPU, low VRAM footprint
What developers value
Data stays on-device — privacy by default
Free & open source, no monthly SaaS fees
Surprising clone quality from a 3-second sample
Uses: voiceover, agent replies, game NPCs, accessibility
Early-stage caveats
v0.x — UI and stability still rough
Generation speed varies with the GPU
Initial model download required
Capable hardware recommended for smooth use
Continue reading The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in ✓ Signed in — this article isn’t included in your current plan.Unlocking the full article…