Dyn-3D: Unveiling and Resolving Ego-Motion Ambiguity in Vision-Language Models
This work introduces Dyn-3D, a benchmark using counterfactual 3D rendering to rigorously decouple visual changes from true kinematic properties and proposes the TempoVista framework, featuring the Kinematic-GSPO algorithm.