A research team from ByteDance and Nanyang Technological University (NTU) has released StoryMem, a framework that generates multi-shot long video stories using an explicit visual memory. It extends pre-trained single-shot video diffusion models into roughly one-minute continuous narrative clips while keeping characters and scenes consistent.
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.