NVIDIA has released Audio2Face-3D as an open-source project, giving developers a collection of models and tools that generate high-fidelity 3D facial animation directly from audio, with accurate lip-sync and audio-driven emotional expression.
September 24, 2025 · NVIDIA ACE
NVIDIA Open-Sources Audio2Face-3D
A freely available toolkit turns a single audio track into high-fidelity 3D facial animation — accurate lip-sync plus audio-driven emotion — now shipped under permissive licensing with production-ready tooling.
60+
frames per second, real-time with CUDA / TensorRT
2
model families — Audio2Face-3D & Audio2Emotion
1
audio source in — full 3D face motion out
The pipeline
From voice to lifelike facial motion
Audio input
a single speech track
→
A2F + A2E models
lip-sync + emotion inference
→
Animation output
mesh · joints · ARKit blendshapes
What ships in the box
Open-source components & licensing
MIT license — SDK plus Maya & Unreal Engine 5 plugins
Apache license — training framework for custom-data retraining
Models — diffusion-based v3.0 & regression-based v2.3, plus Audio2Emotion v2.2 / v3.0
Deployment — NIM Docker containers over gRPC; local & remote inference; multi-stream, CPU fallback
Engines — UE5 plugin v2.5 (UE 5.5 / 5.6), live Maya ACE preview, MetaHuman integration
Why it matters
Lowers the barrier for smaller teams to add audio-driven facial animation without bespoke tooling — a candidate common building block for digital humans.
The catch
Output quality depends on training data; matching a specific character or performance style may require retraining to hold up against tightly controlled commercial pipelines.
Continue reading The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in ✓ Signed in — this article isn’t included in your current plan.Unlocking the full article…