VidParse is presented, an online, training-free framework that treats activity understanding as a graph-constrained inference problem and achieves up to a 10x improvement in complex multi-step parsing accuracy over strong online baselines, all without requiring a single gradient update.
Abstract
Translating continuous, noisy egocentric video streams into discrete, temporally ordered action steps is fraught with visual challenges. Heavy ego-motion, transient occlusions, and the high intra-class variability of unscripted human-object interactions cause standard frame-level online temporal models to struggle, often resulting in severe over-segmentation and structural collapse. To bridge the gap between unstable low-level perception and high-level procedural logic, we present VidParse, an online, training-free framework that treats activity understanding as a graph-constrained inference problem. Rather than relying on learned temporal filters, we dynamically identify semantic transitions using a temporal similarity matrix over manipulation-anchored features, which are extracted from frozen foundation models to prioritize foreground hand-object interactions. A beam search decoder then leverages an induced procedural task graph to explicitly enforce valid action transitions and prune impossible trajectories. By anchoring robust visual segments to hard procedural constraints, our approach preserves long-range state transitions and achieves up to a 10x improvement in complex multi-step parsing accuracy over strong online baselines, all without requiring a single gradient update.
Experiments on egocentric video benchmarks show LogFA significantly improves model generalization to unseen environments while maintaining low computational and data collection costs.
This work introduces a history pathway that enables a vanilla VLA model to summarize observation history into temporally aware latent representations, which captures the evolving 3D world through temporally consistent geometric representations, enabling a deeper understanding of dynamic environments.
Xing-Yu Ding, Yu-Zhong Zhao, Chun-Ming Zhao et al.· 0 citations
Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-grained language-to-motion annotations, and existing predictors either rely on privileged inputs such as video, depth, or CAD models, or rec...
T. Ding, Zhen Luo, Yi-Xuan Yang et al.· 0 citations
Object concepts play a foundational role in human visual cognition, enabling perception, memory, and interaction in the physical world. Inspired by findings in developmental neuroscience - where infants are shown to acquire object understanding through observation of motion - we propose a biologically inspired framewor...
Hao Liang, Xiao-Hui Wang, Zhi-Chao Li et al.· Neural Information Processin...· 0 citations
This work introduces Dyn-3D, a benchmark using counterfactual 3D rendering to rigorously decouple visual changes from true kinematic properties and proposes the TempoVista framework, featuring the Kinematic-GSPO algorithm.
Jiayu Ding, Zhuo-Dong Liu, Lei Zhang et al.· 0 citations
Motion-based tokenization provides a compact representation for event-aligned egocentric gaze streams, while the evaluation identifies how target predictability and event construction shape cross-dataset conclusions.