Vocello (formerly QwenVoice), a local speech-generation tool for Apple Silicon Macs, has shipped as a stable v2.1.0 release supporting macOS 26 and later. It synthesizes speech from text entirely on-device, without routing data through the cloud.
Vocello 2.1.0 · On-Device Speech Generation
Text-to-Speech That Never Leaves Your Mac
Vocello (formerly QwenVoice) turns text into audio on Apple Silicon — cloning a voice from a clip or building one from a plain-language description. Built in Swift on Apple's MLX framework, it runs entirely on the Metal GPU with zero data sent to the cloud.
0
bytes sent to the cloud — fully on-device
10
languages, with automatic detection
9
built-in Custom Voice presets
$0
free, unlimited generation, no subscription
Three ways to make a voice
Custom Voice
Pick from 9 built-in presets — instant and ready to go.
Voice Design
Describe character, age, accent and texture in plain language.
Voice Cloning
Recreate a speaker from a short reference clip, auto-transcribed.
Model download footprint
Faster-than-real-time generation, running even on an 8GB Mac.
The ~7GB download is the main barrier to entry.
STRENGTHS
On-device privacy — nothing goes to the cloud
Voices from plain-language descriptions
Swift-native speed, faster than real time
Strong control for narration, podcasts, audiobooks
LIMITATIONS
~7GB model download is an obstacle
Cloning accuracy depends on clip quality
Quality model memory needs care in big batches
No iOS release yet (iPhone version planned)
System requirements
macOS 26+
Apple Silicon required
Qwen3-TTS + MLX
native Metal GPU inference, no Python
10 styles + intensity
signed & notarized DMG
Continue reading The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in ✓ Signed in — this article isn’t included in your current plan.Unlocking the full article…