This work model step-by-step reasoning as a finite-horizon decision process and introduce Monte Carlo value evaluation on the reasoning tree to provide intermediate supervision signals and effectively enhances the performance of benchmark MLLMs in visual-based spatial understanding and reasoning tasks.
Abstract
Multimodal large language models (MLLMs) have demonstrated significant potential in complex spatial scene understanding and reasoning tasks. However, their open-ended reasoning process is prone to decision errors and error accumulation, leading to instability in answer quality. To address this, we propose an advantage-guided gating framework that dynamically intervenes in and corrects deviations during the reasoning process. Specifically, we model step-by-step reasoning as a finite-horizon decision process and introduce Monte Carlo value evaluation on the reasoning tree to provide intermediate supervision signals. The framework includes Step-Advantage Gate and Trajectory-Advantage Gate, which dynamically select high-value reasoning steps and high-quality complete reasoning trajectories, respectively. During training, we perform supervised learning for the gates using reasoning trees generated via multi-branch sampling, and combine shared-parameter initialization with task-specific heads to achieve cross-task robustness and diversity. During inference, the model greedily selects high-value prefix reasoning steps while choosing the optimal reasoning head based on the problem type, thereby significantly improving the accuracy of the final answer. Furthermore, we constructed the Reasoning-Tree-160k dataset and performed two-stage learning on it. Extensive experiments demonstrate that this advantage-guided gating framework effectively enhances the performance of benchmark MLLMs in visual-based spatial understanding and reasoning tasks. The code is open to the public for research: https://github.com/LingLin-ll/Advantage-Guided-Gate.
This work proposes a new framework that integrates structured reasoning and geometric precision through a teacher-student architecture and outperforms classical reasoning-based baselines in zero-shot reasoning, waypoint accuracy, and inference efficiency.
Vision-language models (VLMs) have advanced scene understanding and enabled explicit reasoning in end-to-end autonomous driving. However, existing methods insufficiently integrate spatial-physical evidence into planning reasoning, while reasoning adaptation remains coarse-grained and falls short of scene-specific planning demands. Furthermore, reasoning-path optimization for higher planning quality remains largely unexplored in autonomous-driving post-training. To address these limitations, we propose FactorDrive, an end-to-end autonomous driving framework for adaptive multi-step reasoning driven by planning-critical factors (PCFs). We first perform large-scale driving-domain instruction tuning to establish foundational driving knowledge. Building on this foundation, we construct PCF-CoT, a chain-of-thought (CoT) dataset that grounds planning reasoning in trajectory-relevant spatial-physical evidence and organizes reasoning around scene-specific PCFs, enabling the composition and depth of reasoning paths to adapt to different planning demands. We further introduce Quality Search-Guided Group Relative Policy Optimization (QS-GRPO), which guides Monte Carlo Tree Search (MCTS) with trajectory-level planning rewards to discover reasoning paths with higher planning quality and uses the resulting responses to optimize the policy through GRPO, thereby improving trajectory planning performance. Extensive experiments on both open-loop (nuScenes) and closed-loop-oriented (NAVSIM) benchmarks demonstrate that FactorDrive achieves state-of-the-art planning performance.
Guolei Huang, Tengfei She, Yuxuan Lu et al.· 0 citations
Generalizing to out-of-distribution scenarios remains a major challenge for traditional robotic manipulation methods trained on closed datasets. Recent approaches leveraging foundation models have significantly improved zero-shot capabilities by utilizing vision-language models, yet many of these methods treat the foundation model as a high-level decision-maker, which limits their adaptability. A key challenge is the gap between high-level decision-making and low-level control. To address the challenge, we propose GAS-Robo, an open-ended manipulation framework that bridges this gap by utilizing a Grid-Action Space (GAS). GAS provides both semantic and spatial information to large language models (LLMs) and enables the direct generation of low-level actions, instead of invoking predefined APIs, thereby enhancing flexibility in trajectory control. The framework comprises two key components: an Environment Filter, which generates a task-aware grid representation of the scene, and an LLM-based Planner, which produces primitive action sequences based on the grid. To improve spatial reasoning and interpretability, a Chain-of-Thought (CoT) mechanism is incorporated into the planner. We benchmark GAS-Robo in the RLBench simulation environment, demonstrating state-of-the-art performance. Furthermore, physical validation on a Franka robotic manipulator platform highlights GAS-Robo's superior generalization across diverse real-world tasks, validating its real-world applicability and robustness in handling diverse manipulation tasks.
Han-Xuan Li, Sen-Wei Xie, Ze-Tao Lin et al.· IEEE Robotics and Automation...· 0 citations
ByDeWay-V2 is proposed, which integrates explicit spatial relational context alongside depth cues, expressed as human-readable predicates that serve as auditable evidence for downstream decision support, showing the framework's suitability for resource-constrained, real-time decision-support settings.
Piyush Jain, Kousik Dasgupta, Rajarshi Roy et al.· arXiv.org· 0 citations
This work designs a structured Chain-of-Thought (CoT) framework that explicitly models 3D environmental perception to ensure robust spatial understanding and reasoning and introduces a novel RL algorithm featuring multi-objective process rewards and a tailored advantage estimation method, facilitating fine-grained credit assignment across distinct segments of the reasoning trajectory.
Zile Zhou, Huining Yuan, Weichen Zhang et al.· 0 citations
TRAM (TRajectory-derived Auxiliary Memory), a training-free method that augments standard decoding with an auxiliary memory pathway derived from the model's own reasoning trajectory, shows that TRAM improves performance over vanilla decoding on mathematical, scientific, and general visual reasoning tasks without additional training.
Kang Liu, Zijing Wang, Yongkang Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.