Skip to content

Rewarding Efficient Reasoning Improves Abstention on Underspecified Tasks in Reasoning Models

Aug 2026 · 0 citations · 50 references
Computer Science

TL;DR

This work takes inspiration from a resource-rational perspective on human cognition and introduces a novel GRPO reward that encourages efficient reasoning about whether the task contains all the information needed to solve it, leading to human-like abstention performance gains.

Abstract

While modern large reasoning models (LRMs) excel at providing correct answers in many tasks, we provide additional evidence for the observation that they often struggle with a critical capability: knowing when to abstain from answering. We analyze this gap by comparing LRM behavior to results from a human study, revealing that human reasoning effort on unanswerable tasks is upper-bounded by answerable tasks, whereas LRMs waste computational resources by generating longer Chains of Thought (CoTs) on unanswerable than on answerable prompts. To overcome this inefficiency, we take inspiration from a resource-rational perspective on human cognition and introduce a novel GRPO reward that encourages efficient reasoning about whether the task contains all the information needed to solve it. Fine-tuning several 4B LRMs with this reward leads to human-like abstention performance gains (+12.8% on average) while retaining answering capabilities and boosting the models'efficiency (44% shorter CoTs on average).

View source

Similar papers

Preprint Aug 2026

BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?

The nature of test-time exploration in RLVR-trained LLMs is investigated by employing controlled maze-solving experiments and extracting a tree structure from mathematical reasoning traces (BODHI-Trees) based on semantic equivalence to delineate between entropy arising from stylistic variations and genuine inferential...

Soumadeep Saha, Krish Sharma, Akshay Chaturvedi et al. · 1 citation
#artificial intelligence Preprint Aug 2026

To Think or Not to Think: Allocating Reasoning Where It Helps

This paper proposes CARE, a method which compares the beneficial length adjustment per question from online sampled responses and applies adaptive length rewards within Group Relative Policy Optimization, with no extra hyperparameters or additional inference cost.

Zheng-Dong He, Yun-Fan Zhou, Jian-Guo Yao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning

Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead. Pruning can reduce this cost, but its effectiveness depends on the calibration data used to estimate parameter importance. Recent work calibrates on the model's own rollouts instead of generic...

Ha Lan Nguyen, Huy Hoang Tran, Trac-Duy Tran et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Jailbreaks for Black-Box Uncertainty Quantification in Large Reasoning Models

While Large Reasoning Models (LRMs) excel at complex reasoning, alignment through reinforcement learning often induces systemic overconfidence. In production environments, where logits may be unavailable, robust black-box uncertainty quantification (UQ) is essential for trustworthiness and safety. Focusing on question-...

Lucas Biechy, Cédric Eichler, Adrien Boiret et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.