Preprint
Sep 2026
Dyn-3D: Unveiling and Resolving Ego-Motion Ambiguity in Vision-Language Models
This work introduces Dyn-3D, a benchmark using counterfactual 3D rendering to rigorously decouple visual changes from true kinematic properties and proposes the TempoVista framework, featuring the Kinematic-GSPO algorithm.
Jiayu Ding, Zhuo-Dong Liu, Lei Zhang et al.
· 0 citations