Preprint
Jul 2026
DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking
DeforM is proposed, a reasoning-guided image-to-video generation framework that directs the model's focus toward physics-critical regions, and introduces a VLM-guided physical reasoning module, DeforM-Reason, to identify target objects and generate spatial-temporal masks.
Yunyi Li, Yu Qiao, Yaohui Wang et al.
· 0 citations