Open access
Mar 2024
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
The ablation study shows that when using SSMs for temporal modeling, incorporating bidirectionality and selective scans enhances video generation performance, and SSM-based models incur lower computational cost to achieve the same Fréchet Video Distance as attention-based models.
Yuta Oshima, Shohei Taniguchi, Masahiro Suzuki et al.
· New generation computing · 13 citations