Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value project...
Yingcheng Liu, Tian-Yi Jiang, Yu-Juan Ding et al.· 0 citations
CUA-Sandbox is introduced, which separates private state capsules from shared runtimes through state-scoped execution and transactional lifecycle operations, including resets and branches, while retaining the original software interfaces and task evaluators.
Xin Yan, Zheng-Bo Jiao, Jia-Qi Liu et al.· 0 citations
WorkDrive is proposed, a framework that constructs perception-grounded causal reasoning for work zones and aligns it with trajectory prediction and achieves progressive improvement over the trajectory-only baseline.
Tianyi Jiang, Wen Zhang, Sihan Yang et al.· arXiv.org· 0 citations
Experiments show that SearchEyes achieves state-of-the-art performance among open-source multimodal search agents, with SearchEyes-27B improving over the strongest open-source baseline by 6.2 points on average.
Zhengbo Jiao, YiMing Cheng, Yilei Jiang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.