A lightweight open-source speech synthesis model called LuxTTS has been released. It supports zero-shot voice cloning that reproduces a speaker's voice from roughly three seconds of reference audio, and is said to generate 48kHz audio at more than 150 times realtime on a single GPU.
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.