A set of custom ComfyUI nodes for the Alibaba Qwen team's open-weight text-to-speech model Qwen3-TTS has been released, bringing voice cloning, natural-language voice design, and support for 10 languages into a node-based workflow.
Open-Source Speech Synthesis · Qwen3-TTS for ComfyUI
Clone a voice from a 15-second clip — locally, in 10 languages
The new FL Qwen3 TTS nodes bring Alibaba's open-source Qwen3-TTS into ComfyUI:
zero-shot voice cloning, natural-language voice design, preset speakers and a fine-tuning dashboard — all running on your own machine.
10
languages supported locally
5–15s
reference clip needed to clone a voice
9
ready-to-use preset speakers
0-shot
cloning — no extra training required
Two model sizes, same family
Parameter count drawn to scale — 1.7B is roughly 2.8× the lightweight 0.6B build.
Voice Cloning
Reproduce a speaker from a short reference clip — with Whisper auto-transcribing the reference text.
Voice Design
Create custom voices from plain language, e.g. “warm British female voice.”
Fine-tuning UI
A real-time dashboard for training — plus 9 preset speakers to start instantly.
Full cloning workflow in ComfyUI
Reference clip
3–15 sec audio
→
ASR transcribe
Whisper / Qwen3-ASR
→
Qwen3-TTS
generate speech
→
Video / avatar
audiobook, dialogue
What developers value
High-quality clones from short samples
Whisper auto-generates reference text
Strong fit with video-generation workflows
Practical limitations
Clips over ~30s can hang generation
Missing ref_text lowers similarity
Training wants 32GB+ RAM, 12GB+ GPU
Continue reading The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in ✓ Signed in — this article isn’t included in your current plan.Unlocking the full article…