Skip to content

Nudgeability: Reasoning Models Follow Confidence Signals Without Tracking Their Own Competence

Sep 2026 · 0 citations · 24 references
Computer Science

TL;DR

Nudgeability offers a simple, post-training-free way to evaluate both sensitivity and targeting as endogenous self-reflection mechanisms mature as endogenous self-reflection mechanisms mature.

Abstract

Reasoning language models that can call tools must decide during inference whether to answer unaided or delegate. Any self-reflection mechanism for this must answer three questions: where the reflective signal comes from (verbal reports, output distributions, hidden states, a separate predictor), how it is presented to the model (numerical prediction, confidence token, prompt injection), and whether it changes the model's subsequent action. We isolate the third question. At a fixed point in otherwise identical reasoning trajectories, we insert a single first-person sentence expressing either confidence or doubt; the model then continues reasoning and chooses whether to answer directly or call a tool. Comparing these counterfactual continuations measures the causal effect of the reflective signal on delegation. We call this behavioral response Nudgeability and measure it along two dimensions: sensitivity, how strongly confidence and doubt change delegation rates, and targeting, whether delegation increases for problems the model cannot solve unaided and decreases for those it can. Across nine small-to-medium open-weight reasoning models from three families (Qwen, Gemma, and GLM) and two tasks, models are consistently sensitive: doubt increases delegation and confidence decreases it, with a median confidence-to-doubt swing of 20.6 percentage points, and 53 to 70 points for the larger provider-served models. This responsiveness is poorly targeted: a median 42% of induced flips are well-targeted, only a +2 percentage-point lift over a random-selection baseline. Confidence language is thus a strong control surface for delegation, but current models use it only weakly in accordance with their actual competence. Nudgeability offers a simple, post-training-free way to evaluate both sensitivity and targeting as endogenous self-reflection mechanisms mature.

View source

Similar papers

#natural language process... Preprint Sep 2026

When Confidence Rises Too Early: Detecting Shortcut Reasoning via Premature Answer Commitment

The reasoning trajectory of a Large Language Model (LLM) is often treated as a verbalized description of its internal reasoning. However, such trajectories can be unfaithful: a model may rely on shortcuts to reach an answer and then post-rationalize the decision with a seemingly coherent chain of thought. Detecting thi...

Zhao-Han Zhang, Jun-Jie Liu, Chengzhengxu Li et al. · 0 citations
#natural language process... Preprint Sep 2026

Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models

Reinforcement learning with binary correctness rewards trains correctness, not calibrated confidence. The confidence that reasoning models verbalize is systematically overconfident, and the problem is not merely one of scale: verbalized confidence tracks how willing a model is to commit to an answer, not how likely the...

Ya-Dong Xi, Rong-Sheng Zhang, Tang-Jie Lv et al. · 0 citations
#artificial intelligence Preprint Aug 2026

DERELAB: Probing Defeasible Reasoning and Confirmation Bias in LLMs with a Generative Benchmark

DeReLab is introduced, a generative framework that produces multi-turn belief-updating conversations from parameterized graph structures across default and inheritance reasoning, with formally verified ground truth at every turn, enabling controlled measurement of how models respond to confirming and disconfirming evid...

Jayanta Sadhu, S. Shahad, Kenneth Marino · 1 citation
#artificial intelligence Preprint Sep 2026

REALHOP: Rethinking Multi-Hop Reasoning Evaluation via Behavioral Auditing

Complex questions often require multi-hop reasoning that connects facts distributed across sources or distant regions of a long context through intermediate steps. Benchmarks commonly evaluate this ability with questions built around predefined reasoning chains, treating a correct answer as evidence that the intended c...

Ji-Hua Tao, Xiao-Kun Yuan, Yao-Ming Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Are Stated Reasoning Steps Causally Load-Bearing?

This work uses synthetic multi-hop lookup tasks to measure faithfulness causally at the activation level, specifically on self-generated reasoning, and aims to measure faithfulness causally at the activation level, specifically on self-generated reasoning.

Abhiram Bhupatiraju, Rayan Nyaupane · 0 citations
#artificial intelligence Conference Open access Sep 2026

Self-Reports Are Not Verification

An environment-grounded audit is introduced in which every intermediate proposal receives an exact outcome in an evolutionary Contexto search whose feedback function assigns every valid guess an exact rank without human annotation.

En-Rong Pan, Ryan Zhou, Ting Hu · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.