A lightweight Vietnamese-focused text-to-speech model, "ValtecTTS," has been released and updated on GitHub. With 74.8M parameters and GPU-free CPU operation, it supports zero-shot voice cloning that reproduces a speaker's voice from just 3–10 seconds of reference audio.
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.