This work introduces \lifd{} (Look, Imagine, Focus, and Do), a framework for persistent, 3D-aware scene memory that learns scene tokens through multi-view agreement, then completes them from a single RGB view and recurrent memory using rectified flow.
Wen-Bo Li, Yi-Teng Chen, Wen-Hao Li et al.· 0 citations
Learned world models are becoming essential to general-purpose agents: by predicting action consequences, they support planning and decision-making while reducing reliance on costly trial and error. This reliance creates a fundamental ambiguity: when a world-model-guided decision fails, the trajectory alone may not rev...
Hua-Tai Zhu, Qiang Chen, Zi-Qian Kou et al.· 0 citations
StateTrace is proposed, a novel object-centric framework that endows VideoLLMs with an explicit mechanism for hidden state reasoning in long videos, and builds a reusable spatiotemporal state memory that organizes object trajectories, inter-object relations, and state-transition events into a structured reasoning subst...
Trajectory-Aware Commit Gating (TACG), a training-free gate-level decoder that anchors token identities to the base posterior and uses trajectory-aware signals only to decide whether the current proposal is ready to commit, is proposed.
Chengcheng Wang, Tingzhang Luo, Wenhao Li et al.· arXiv.org· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.