Skip to content

Do Frontier Models Seek Safety Evidence Before Acting?

Sep 2026 · 0 citations · 10 references
Computer Science

TL;DR

SAFE, a controlled benchmark in which models make deployment decisions with optional evidence that varies in retrieval cost, probability, severity, and presentation, is introduced and suggests that deployment-time safety depends not only on how models respond to known risks, but also on whether they acquire the evidence needed to know that acting is safe.

Abstract

Frontier models are often evaluated on how they respond to safety information once it is already in context. We study an earlier decision point: whether models choose to acquire safety-relevant evidence before acting. We introduce SAFE, a controlled benchmark in which models make deployment decisions with optional evidence that varies in retrieval cost, probability, severity, and presentation. Across GPT-5.5, o3, Claude Opus 4.8, and Claude Sonnet 4.6, we find distinct evidence-acquisition policies: Opus inspects nearly by default, o3 is the most skip-heavy and threshold-sensitive, and GPT-5.5 and Sonnet occupy intermediate regimes. Inspection increases strongly with severity and decreases with retrieval cost, whereas probability has much weaker behavioral influence: increasing the stated likelihood of a problem from 10% to 70% changes inspection by at most 21 percentage points. Despite these differences, Stage 1 rationales are dominated by expected-value reasoning across models. A cost-obligation decomposition further shows that avoidance is driven primarily by retrieval friction and explicit threats to the deployment payoff rather than by the remediation duties created by knowing. Counterfactual interventions reveal a further mismatch between behavior and explanation: evidence framing can strongly change decisions near the inspection boundary while going largely unmentioned, whereas probability is frequently cited despite having little causal influence. These results suggest that deployment-time safety depends not only on how models respond to known risks, but also on whether they acquire the evidence needed to know that acting is safe.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Dual-Frontier: When Can an Agent Trust Its World Model?

Learned world models are becoming essential to general-purpose agents: by predicting action consequences, they support planning and decision-making while reducing reliance on costly trial and error. This reliance creates a fundamental ambiguity: when a world-model-guided decision fails, the trajectory alone may not rev...

Hua-Tai Zhu, Qiang Chen, Zi-Qian Kou et al. · 0 citations
Review Open access Sep 2026

On the Cost of Conforming to Reviewers’ Expectations and the Potential Benefit of Prediction Competitions

Although academic peer review offers many important benefits, it can also impede scientific exploration. For instance, when reviewers share restrictive working assumptions, researchers may be incentivized to conform to them, even when alternative conjectures could better advance scientific understanding. We suggest mit...

Ido Erev, Adi Tarabeih, Rachel Barkan · 0 citations
Preprint Aug 2026

Bayesian Expected Uncertainty Reduction (B-EUR) Model: A Computational Account of What Makes Design Options Worth Trying

This paper proposes the Bayesian Expected Uncertainty Reduction (B-EUR) model, which formalizes the value of trying a candidate design action as its expected reduction of epistemic uncertainty about action--outcome relations. The model addresses one part of the Uncertainty Driven Action (UDA) model's open question conc...

Shimon Honda, Takuma Miyaguchi, Koji Koizumi et al. · 0 citations
Preprint Aug 2026

Effort without Evidence

Advice can determine not only how past evidence is interpreted but whether new evidence will be produced. I study a sequence of decision makers who observe public success or failure but not one another's effort, and who pass costless causal assessments to their successors. Because a sender values success while underwei...

G. Lukyanov · 1 citation
Preprint Aug 2026

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings

Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations study explicit-influence settings: setups where the prompt directly incentivizes the model to hide something, e.g., by instructing it to perform a hidden side-task. A com...

Agatha Duzan, Asa Cooper Stickland · 2 citations
#artificial intelligence Preprint Sep 2026

Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding

Language-model agents increasingly answer questions over customer-relationship management (CRM) records, such as whether to qualify a sales lead. We identify a failure mode not addressed by a stronger model: when the context contains an assertion by a party with an incentive toward optimism - here the sales representat...

Rahul Balakavi · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.