This work takes inspiration from a resource-rational perspective on human cognition and introduces a novel GRPO reward that encourages efficient reasoning about whether the task contains all the information needed to solve it, leading to human-like abstention performance gains.
Abstract
While modern large reasoning models (LRMs) excel at providing correct answers in many tasks, we provide additional evidence for the observation that they often struggle with a critical capability: knowing when to abstain from answering. We analyze this gap by comparing LRM behavior to results from a human study, revealing that human reasoning effort on unanswerable tasks is upper-bounded by answerable tasks, whereas LRMs waste computational resources by generating longer Chains of Thought (CoTs) on unanswerable than on answerable prompts. To overcome this inefficiency, we take inspiration from a resource-rational perspective on human cognition and introduce a novel GRPO reward that encourages efficient reasoning about whether the task contains all the information needed to solve it. Fine-tuning several 4B LRMs with this reward leads to human-like abstention performance gains (+12.8% on average) while retaining answering capabilities and boosting the models'efficiency (44% shorter CoTs on average).
A formal analysis showing that joint optimization of the two objectives induces gradient conflict in early training, motivating the sequential design of ERR+, a two-phase RLVR framework grounded in this observation.
Xinle Jiang, Min-Hao Wang, Wen Wu et al.· 0 citations
The nature of test-time exploration in RLVR-trained LLMs is investigated by employing controlled maze-solving experiments and extracting a tree structure from mathematical reasoning traces (BODHI-Trees) based on semantic equivalence to delineate between entropy arising from stylistic variations and genuine inferential...
Soumadeep Saha, Krish Sharma, Akshay Chaturvedi et al.· 1 citation
This paper proposes CARE, a method which compares the beneficial length adjustment per question from online sampled responses and applies adaptive length rewards within Group Relative Policy Optimization, with no extra hyperparameters or additional inference cost.
Zheng-Dong He, Yun-Fan Zhou, Jian-Guo Yao et al.· 0 citations
Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead. Pruning can reduce this cost, but its effectiveness depends on the calibration data used to estimate parameter importance. Recent work calibrates on the model's own rollouts instead of generic...
Ha Lan Nguyen, Huy Hoang Tran, Trac-Duy Tran et al.· 0 citations
LightTIR, a dual-penalty reward framework, is proposed to achieve efficient TIR and can reduce redundancy and trajectory expansion while maintaining answer correctness, achieving more efficient RL-based TIR.
Yichen Xiao, Siyu Gong, Li-Nan Yue· Proceedings of the 32nd ACM...· 0 citations
While Large Reasoning Models (LRMs) excel at complex reasoning, alignment through reinforcement learning often induces systemic overconfidence. In production environments, where logits may be unavailable, robust black-box uncertainty quantification (UQ) is essential for trustworthiness and safety. Focusing on question-...
Lucas Biechy, Cédric Eichler, Adrien Boiret et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.