This work proposes B rain-inspired Supervised Supervised Reflection (BUS), a label-free training framework to enhance reflective reasoning capability in challenging image analysis and validate that backward prediction capability is critical for VLM reasoning.
TRAM (TRajectory-derived Auxiliary Memory), a training-free method that augments standard decoding with an auxiliary memory pathway derived from the model's own reasoning trajectory, shows that TRAM improves performance over vanilla decoding on mathematical, scientific, and general visual reasoning tasks without additi...
Kang Liu, Zi-Jing Wang, Yongkang Liu et al.· 0 citations
Multimodal Large Language Models (MLLMs) often struggle with complex mathematical visual reasoning primarily due to a lack of fine-grained perception, causing initial visual hallucinations to directly trigger cascading reasoning failures. In traditional end-to-end reinforcement learning (RL), sparse rewards fail to dec...
Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based methods have improved navigation, they often preserve excessive task-irrelevant detail, weakening ge...
Yihao Wu, Chen-Yi Xu, Li-Qi Yan et al.· arXiv.org· 0 citations
A self-supervised framework that trains models on systematically simplified versions of abstract reasoning tasks containing incomplete but structured concept cues enables models to form internal abstractions under limited resources and later apply them to more complex problems.
Lingxiao Yang, Muyang Lyu, Ying-Jie Wang et al.· Science Advances· 0 citations
This work proposes BrainWAM, a structured action-space coordination framework that converts semantic reasoning and predictive world modeling into two specialized action-oriented pathways, and aligns them at the level of compact action representations, and introduces an asynchronous rectified-flow inference strategy wit...
Bing Zhan, Shuyao Shang, Shuo Lu et al.· 1 citation
A novel training-free inference strategy for MLLMs that explicitly decouples perception and reasoning is presented, and a novel metric, the vision-to-text attention ratio, is proposed, to dynamically gauge the model's cognitive focus.
Haoqiang Kang, Liupeng Li, Kuo-Feng Gao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.