ainewsblitz.com

Breaking

Alibaba unveils Wan-Streamer, a single model generating synced speech and video

  • Foundation Models
  • Media Generation
  • Research & Papers

The Wan team led by Alibaba's Tongyi Lab has introduced Wan-Streamer, a duplex interaction model that listens, watches, understands and speaks through a single Transformer, generating synchronized speech and video at once. The latest v0.2 was posted to arXiv in July 2026, with a paper released alongside a dedicated site, wan-streamer.com. Its predecessor v0.1 appeared on arXiv in June the same year, positioning the work as a continuing research line.

Continue reading

The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.

$20
Read this article
$29/month
Unlimited — all 8,014 articles, the full archive, and comprehension quizzes
Save 72%
$98/year
≈ $8.17/month
Unlimited, billed once a year