This work introduces Dyn-3D, a benchmark using counterfactual 3D rendering to rigorously decouple visual changes from true kinematic properties and proposes the TempoVista framework, featuring the Kinematic-GSPO algorithm.
Abstract
As Vision-Language Models (VLMs) tackle dynamic 3D spatial reasoning, ego-motion perception becomes essential to resolve monocular scale ambiguity. However, current models often overfit to smooth trajectory priors rather than genuinely understanding physical motion. Consequently, their spatial reasoning degrades severely under large displacements, a phenomenon we term Kinematic Collapse. This failure stems from spurious visual-motion correlations in natural videos and a lack of explicit physical supervision. To evaluate this, we introduce Dyn-3D, a benchmark using counterfactual 3D rendering to rigorously decouple visual changes from true kinematic properties. Furthermore, we propose the TempoVista framework, featuring the Kinematic-GSPO algorithm. By embedding metric physical ground truth into policy optimization, TempoVista explicitly grounds visual representations in 3D space. Experiments demonstrate that our approach significantly improves both motion estimation and robust spatial reasoning by utilizing camera dynamics as an effective geometric calibration signal.
Recent visual generators produce high-fidelity images yet often violate physical consistency under ego-motion, limiting their use for spatial reasoning and embodied planning. Existing benchmarks largely focus on isolated images or single-step quality, leaving this challenge underexplored. We introduce EgoGenEval, a geo...
Yi-Lin Long, Chen-Ming Zhu, Zi-Tang Gou et al.· 0 citations
This work introduces a history pathway that enables a vanilla VLA model to summarize observation history into temporally aware latent representations, which captures the evolving 3D world through temporally consistent geometric representations, enabling a deeper understanding of dynamic environments.
Xing-Yu Ding, Yu-Zhong Zhao, Chun-Ming Zhao et al.· 0 citations
GaussianWAM is proposed, a training-time representation-enhancement framework that organizes geometric and semantic supervision through a 3D Gaussian field and improves performance on standard LIBERO and shows positive transfer trends on RoboTwin and real-world manipulation.
Zi-Jian Zhang, Yu-Qing Jiang, Wei-Tao Zhou et al.· 1 citation
UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter that converts RGB-derived bidirectional flow into motion-canonical visual evidence in three stages. F...
Jing-Pu Yang, Feng-Xian Ji, Ming-Xuan Cui et al.· 0 citations
Experiments show that RESELF outperforms state-of-the-art methods designed for the individual tasks across depth estimation, camera tracking, and full-body motion estimation.
Kai Guan, Minchao Jiang, Ruichen WangLi et al.· 0 citations
VidParse is presented, an online, training-free framework that treats activity understanding as a graph-constrained inference problem and achieves up to a 10x improvement in complex multi-step parsing accuracy over strong online baselines, all without requiring a single gradient update.
Anubhav Gupta, A. Kambhamettu, Vatsal Agarwal et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.