Skip to content
Preprint

DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking

Jul 2026 · 0 citations · 35 references
Computer Science

TL;DR

DeforM is proposed, a reasoning-guided image-to-video generation framework that directs the model's focus toward physics-critical regions, and introduces a VLM-guided physical reasoning module, DeforM-Reason, to identify target objects and generate spatial-temporal masks.

Abstract

Video generation models achieve high visual quality but often struggle to generate physics-aware videos. Unlike rigid-body motion, which can be described by explicit trajectories or formulas, complex deformation dynamics remain challenging to synthesize. We observe that a lack of physical reasoning for localizing dynamic areas allows irrelevant regions to dilute the model's attention, leading to generation failure. In this paper, we propose DeforM, a reasoning-guided image-to-video generation framework that directs the model's focus toward physics-critical regions. To reason about and localize these critical regions, we introduce a VLM-guided physical reasoning module, DeforM-Reason, to identify target objects and generate spatial-temporal masks. For physical guidance, we develop two alternative strategies: DeforM-Free for training-free mechanism analysis and DeforM-Injection as a powerful training-based generator. Experimental results demonstrate that DeforM improves the realism of generated deformation scenarios, outperforming baseline models in both visual quality and physical consistency.

View source

Similar papers

Preprint Jul 2026

VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation

Experiments on an unseen validation set show that VIPER achieves stronger reference-video physical similarity and higher human preference than representative video generation and video-as-prompt baselines, while maintaining competitive general video quality.

Tianxi Chen, Hanmo Chen, Huajin Chen et al. · 0 citations
Preprint Jul 2026

Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

PhyParam is presented, a physics-guided image-to-video diffusion model that conditions on object-level forces, masses, friction, restitution, and scene-level gravity via a lightweight physical-attention routing mechanism, and further improves motion learning with semantic-structural feature-space supervision.

Yan-Xun Li, Hao Wen, Bingze Song et al. · 0 citations
#generative ai Preprint Aug 2026

MotionPhys: Detecting AI-Generated Videos via Physical Consistency of Optical-Flow Trajectories

Modern AI video generation models can produce videos with high visual fidelity and seemingly smooth temporal transitions. However, visual realism does not necessarily imply physical motion consistency. Existing generative models mainly optimize distribution matching in pixel or latent spaces, without explicitly enforcing real-world constraints such as inertia, continuous forces, and trajectory geometry. Our experiments show that AI-generated videos remain visually plausible over short sequences of consecutive frames, yet fail to preserve physical motion consistency throughout a complete object action, resulting in systematic statistical discrepancies in their motion trajectories. Based on this observation, we introduce MotionPhys, a lightweight and interpretable framework that treats sparse motion trajectories as physical evidence rather than relying on appearance artifacts or generator-specific traces. By modeling the geometric evolution of trajectories across multiple temporal scales, MotionPhys reveals subtle motion inconsistencies that are difficult to capture with conventional visual cues and transforms them into a compact representation for efficient detection. Experiments on multiple datasets show that MotionPhys can effectively detect physical inconsistencies in generated videos and generalizes well across different video generators.

Hao He, Hao Tan, Zichang Tan et al. · 0 citations
Jul 2026

Motion-driven 4D scene generation

Guo-Wei Yang, Qun-Ce Xu, Zhao Wei et al. · 0 citations
2025

PhysDiff-VTON: Cross-Domain Physics Modeling and Trajectory Optimization for Virtual Try-On

We present PhysDiff-VTON, a diffusion-based framework for image-based virtual try-on that systematically addresses the dual challenges of garment deformation modeling and high-frequency detail preservation. The core innovation lies in integrating physics-inspired mechanisms into the diffusion process: a pose-guided deformable warping module simulates fabric dynamics by predicting spatial offsets conditioned on human pose semantics, while wavelet-enhanced feature decomposition explicitly preserves texture fidelity through frequency-aware attention. Further enhancing generation quality, a novel sampling strategy optimizes the de-noising trajectory via least action principles, enforcing temporal coherence, spatial smoothness, and multi-scale structural consistency. Comprehensive evaluations across multiple datasets demonstrate significant improvements in both geometric plausibility and perceptual quality compared to existing approaches. The framework establishes a new paradigm for synthesizing photorealistic try-on images that adhere to physical constraints while maintaining intricate garment details, advancing the practical applicability of diffusion models in fashion technology.

Shibin Mei, Bingbing Ni · 1 citation
Preprint Aug 2026

RigidBench: Evaluating Rigid-Body Physics in Video Generation Models

RididBench is introduced, a simulator-grounded benchmark that compares a generated continuation with a reference rollout from the same initial frame and motion description, with per-frame masks, depth, 6-DoF trajectories, and contacts available for scoring.

Swarnim Jain, Shangzhe Wu · 2 citations