Skip to content
Preprint

Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

Aug 2026 · 1 citation
Computer Science

TL;DR

It is found that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.

Abstract

Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. We quantify this mismatch with Behavioral Lift, a metric that measures how much correctness changes when a behavior is present versus absent in a model's reasoning trace. Across 15 models and 6 benchmarks spanning text-only and vision-language reasoning, we annotate 15,282 traces with a taxonomy whose core behaviors are defined for both LLM and VLM traces. We find evidence for an Amplification-Lift Gap, in which thinking models strongly amplify self-correction, hypothesis testing, and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment, and self-awareness. Confidence calibration is among the strongest positive signals of correctness in both modalities, yet is barely amplified; uncertainty acknowledgment is amplified by 3--7$\times$, yet is weakly or negatively associated with correctness. We find that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Thinking effort aligns between humans and reasoning models in abductive reasoning

This work isolates abductive reasoning from deductive reasoning: unlike deductive tasks, its difficulty cannot be inferred from formal structure and offers no shortcuts a model could exploit to mimic effort without genuine search, providing firmer ground for empirical claims of shared effort.

Henry Arthur · 0 citations
#machine learning Preprint Sep 2026

Do Reasoning Representations Help Humans Evaluate LLM Outputs?

This work conducts a controlled human study of six reasoning formats across tasks of varying complexity, supported by a web-based framework that randomizes task domains, problem instances, and representation order.

Jaewoo Lim, Sungbok Shin, San Hong · 0 citations
Preprint Aug 2026

Cognitive Profiling of LRMs'Reasoning Traces Using Bloom's Taxonomy

This work introduces a framework for automatic annotation of reasoning steps through the lens of Bloom's Taxonomy, which classifies thinking into six cognitive levels, such as Remembering, Applying and Evaluating, and demonstrates that thinking-type information derived from reasoning traces correlates with correctness,...

Maria-Eleni Zoumpoulidi, Georgios Paraskevopoulos, A. Potamianos · 0 citations
#machine learning Preprint Sep 2026

When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models

This work studies logical verification in five open-weight transformer models using matched valid--invalid premise--claim pairs that vary across inference families, semantic domains, templates, and inference families, suggesting that representing validity, expressing it in behavior, and using it causally are distinct.

Smitha Muthya Sudheendra, Jaideep Srivastava · 1 citation
Preprint Aug 2026

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although these modes differ in reasoning budget, their delivered responses should satisfy the same user-facing standard. Correctness alone may not charac...

Xin-Ming Wang, Wei-Nong Wang, Hong-Ming Yang et al. · 2 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.