Skip to content
Preprint

Dyn-3D: Unveiling and Resolving Ego-Motion Ambiguity in Vision-Language Models

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

This work introduces Dyn-3D, a benchmark using counterfactual 3D rendering to rigorously decouple visual changes from true kinematic properties and proposes the TempoVista framework, featuring the Kinematic-GSPO algorithm.

Abstract

As Vision-Language Models (VLMs) tackle dynamic 3D spatial reasoning, ego-motion perception becomes essential to resolve monocular scale ambiguity. However, current models often overfit to smooth trajectory priors rather than genuinely understanding physical motion. Consequently, their spatial reasoning degrades severely under large displacements, a phenomenon we term Kinematic Collapse. This failure stems from spurious visual-motion correlations in natural videos and a lack of explicit physical supervision. To evaluate this, we introduce Dyn-3D, a benchmark using counterfactual 3D rendering to rigorously decouple visual changes from true kinematic properties. Furthermore, we propose the TempoVista framework, featuring the Kinematic-GSPO algorithm. By embedding metric physical ground truth into policy optimization, TempoVista explicitly grounds visual representations in 3D space. Experiments demonstrate that our approach significantly improves both motion estimation and robust spatial reasoning by utilizing camera dynamics as an effective geometric calibration signal.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Beyond Visual Quality: Evaluating Physical Consistency under Ego-Motion with EgoGenEval

Recent visual generators produce high-fidelity images yet often violate physical consistency under ego-motion, limiting their use for spatial reasoning and embodied planning. Existing benchmarks largely focus on isolated images or single-step quality, leaving this challenge underexplored. We introduce EgoGenEval, a geo...

Yi-Lin Long, Chen-Ming Zhu, Zi-Tang Gou et al. · 0 citations
Preprint Aug 2026

Temporal Forcing: 4D Representation Alignment for Vision-Language-Action Models

This work introduces a history pathway that enables a vanilla VLA model to summarize observation history into temporally aware latent representations, which captures the evolving 3D world through temporally consistent geometric representations, enabling a deeper understanding of dynamic environments.

Xing-Yu Ding, Yu-Zhong Zhao, Chun-Ming Zhao et al. · 0 citations
Preprint Aug 2026

GaussianWAM: Distilling Geometry and Semantics from 3D Gaussian Fields into World-Action Models

GaussianWAM is proposed, a training-time representation-enhancement framework that organizes geometric and semantic supervision through a 3D Gaussian field and improves performance on standard LIBERO and shows positive transfer trends on RoboTwin and real-world manipulation.

Zi-Jian Zhang, Yu-Qing Jiang, Wei-Tao Zhou et al. · 1 citation
Preprint Sep 2026

EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning

UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter that converts RGB-derived bidirectional flow into motion-canonical visual evidence in three stages. F...

Jing-Pu Yang, Feng-Xian Ji, Ming-Xuan Cui et al. · 0 citations
Preprint Sep 2026

Seeing the World and the Self from Egocentric Video

Experiments show that RESELF outperforms state-of-the-art methods designed for the individual tasks across depth estimation, camera tracking, and full-body motion estimation.

Kai Guan, Minchao Jiang, Ruichen WangLi et al. · 0 citations
Preprint Aug 2026

VidParse: Online Parsing of Egocentric Procedures Like a Pro

VidParse is presented, an online, training-free framework that treats activity understanding as a graph-constrained inference problem and achieves up to a 10x improvement in complex multi-step parsing accuracy over strong online baselines, all without requiring a single gradient update.

Anubhav Gupta, A. Kambhamettu, Vatsal Agarwal et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.