BREAKING
Alibaba unveils Wan-Streamer model
One Transformer listens, watches, speaks
Replacing the cascade pipeline
1
VAD + ASR
↓
2
LLM
↓
3
TTS
↓
4
Video gen
0
ms
Model latency
0
ms
Total latency
0
ms
Streaming unit
v0.2 sharpens the output
v0.1 width
192
v0.2 width
640
v0.2 fps
25
Open weights still undecided
AI NEWS BLITZ
Alibaba's Tongyi Lab has revealed Wan-Streamer, a model generating synced speech and video.