R-Qwen is a recursive reasoning framework built upon a pretrained Qwen backbone, which consistently outperforms prior recursive reasoning models and substantially larger LLMs while using a comparable number of trainable parameters, suggesting that recursive reasoning mechanisms and pretrained language model priors are complementary approaches for improving symbolic puzzle-solving.
Abstract
Hard symbolic-reasoning tasks such as Sudoku, maze pathfinding, and ARC remain challenging for LLMs due to their fixed-depth autoregressive reasoning, which limits systematic search, refinement, and backtracking. While recursive models such as Hierarchical Reasoning Model (HRM) and Tiny Recursive Model (TRM) address this limitation through iterative latent-state refinement, they are typically task-specific and do not leverage pretrained language priors. We propose R-Qwen, a recursive reasoning framework built upon a pretrained Qwen backbone. R-Qwen repeatedly refines a candidate solution through programmatic self-recursion and deep supervision, combining the structured iterative computation of recursive models with the linguistic and reasoning priors of pretrained LLMs. We further adapt Hierarchical Supervision Weighting (HSW) to autoregressive models by exponentially weighting losses across recursive steps. HSW reduces gradient variance by at least 50\%, improves the signal-to-noise ratio of stochastic gradients, and accelerates convergence. Across eight challenging benchmarks, R-Qwen consistently outperforms prior recursive reasoning models and substantially larger LLMs while using a comparable number of trainable parameters. Notably, on ARC-AGI dataset, our model achieves a 27.6\% improvement over the baseline, highlighting the effectiveness of recursive refinement for general symbolic reasoning. These results suggest that recursive reasoning mechanisms and pretrained language model priors are complementary approaches for improving symbolic puzzle-solving. Code and models will be released after acceptance.
This work proposes a novel test-time alignment approach that leverages trajectory-guided structured sampling for dynamic inference-time refinement, achieving better alignment with visual grounding and ensuring logical consistency, and demonstrates that this approach significantly improves accuracy without incurring pro...
Tian-Bao Jiang, Wei-Cong Ni, Gerard de Melo et al.· 1 citation
DLMR is a parameter-efficient mechanism that equips MLLMs with Dual Latent Memories: a visual memory that compresses image evidence and a reasoning memory that tracks intermediate conclusions and constraints with only a small number of additional trainable parameters.
Hao-Xuan Ma, Jin-Fei Qi, Yi-Cheng Xiao et al.· 0 citations
This work introduces SPEAR (Symbolic Process Evaluation and Alignment Reward), a training-free and plug-and-play process reward method for sequence-level on-policy distillation that effectively bridges the reasoning gap between student and teacher models via sequence-level distillation with efficient dense process rewa...
Zhuo-Chun Li, Yuelyu Ji, Yiming Zeng et al.· 0 citations
This work introduces StructReward, a compute-efficient framework that provides dense reinforcement signals through structured step-level reward alignment and substantially reduces the computational overhead of multimodal reinforcement learning.
This work model step-by-step reasoning as a finite-horizon decision process and introduce Monte Carlo value evaluation on the reasoning tree to provide intermediate supervision signals and effectively enhances the performance of benchmark MLLMs in visual-based spatial understanding and reasoning tasks.
Ling Lin, Yang Bai, Cong-Cong Zhu et al.· 0 citations
Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual information. We show...
Ahmadreza Jeddi, En-Ming Zhang, Jasper Gerigk et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.