Skip to content
Preprint

Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable

Aug 2026 · 0 citations · 5 references
Computer Science

Abstract

An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data. What unlocks confident action is not information but the authority of its packaging. The failure is narrow and locatable. Incapacity is not the answer: on matched answerable questions attached to the same panels, the same models answer essentially always, at near-perfect accuracy. Nor is it belief - stated probabilities barely move across the gradient that swings action by 48 points, and score worse than a climatological baseline. Missing judgment isn't it either: asked to classify a question's knowability before acting, models call it irreducible 90% of the time and then commit on just 0.4% of those. The act/don't-act gate is what fails, and the effect is concentrated in a few models rather than universal. Because the gate is separable, it can be trained. Supervised fine-tuning of a 3B model on 540 synthetic cases, predominantly dice, coins, jars and timers, drives commitment to 0.0% on the original cases and transfers to three unseen domains. It does not survive everything: the gate holds exactly when the response format leaves room to reason, and rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly. The gate is trainable and context-fragile, and deployment needs both halves of that sentence.

View source

Similar papers

Preprint Aug 2026

Confident at the moment of action: belief miscalibration in LLM play under hidden information

This work tests a hidden-information chess variant where royal status can be secretly, repeatedly relocated between pieces, and where an agent's stated probability distribution over the opponent's hidden royal piece -- elicited every turn, separately from the move it chooses -- is scored against ground truth recoverabl...

Bhushan Kashinath Joshi · 1 citation
Case report Open access Sep 2026

Is It a Lie If I Don’t Know? Mechanisms and Mitigation of Dishonesty Under Ignorance

Ignorance of facts and laws may provide an excuse for self-serving reporting behavior, even at the risk of telling the untruth. This paper examines what decision-makers report when they do not know their true entitlement to a financial gain, why they do so, and how the resulting dilemma under ignorance can be mitigated...

Sven A. Simon · 0 citations
#artificial intelligence Preprint Sep 2026

The Missing"I Don't Know": Why Three Reasoning-Reliability Findings Converge on Calibrated Abstention

It is argued these findings converge on a single intervention: calibrated abstention is what each independently identifies as the missing capability, even though the unavailability they document, a capability gap, a policy gap, and a recursion-theoretic gap, has a different source in each case.

Srijith Ravikumar · 0 citations
Preprint Aug 2026

Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives

A knowledge-verified benchmark that first confirms through a neutral probe that an agent knows a user's entitlement, and then evaluates whether it makes false claims once an incentive to deny that entitlement is introduced, which reduces the confound between lying and not knowing and enables more rigorous auditing and...

Zhe-Yuan Liu, Wei-Liang Zhao, Xiangchi Yuan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Do Frontier Models Seek Safety Evidence Before Acting?

SAFE, a controlled benchmark in which models make deployment decisions with optional evidence that varies in retrieval cost, probability, severity, and presentation, is introduced and suggests that deployment-time safety depends not only on how models respond to known risks, but also on whether they acquire the evidenc...

Omer Tafveez · 0 citations
Jul 2026

What AI Red-Team Evaluations Can and Cannot Prove

This work defines the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and uses it to locate that boundary exactly.

Bandana Kaur · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.