Skip to content

Do Frontier Models Seek Safety Evidence Before Acting?

Sep 2026 · 0 citations · 10 references
Computer Science

TL;DR

SAFE, a controlled benchmark in which models make deployment decisions with optional evidence that varies in retrieval cost, probability, severity, and presentation, is introduced and suggests that deployment-time safety depends not only on how models respond to known risks, but also on whether they acquire the evidence needed to know that acting is safe.

Abstract

Frontier models are often evaluated on how they respond to safety information once it is already in context. We study an earlier decision point: whether models choose to acquire safety-relevant evidence before acting. We introduce SAFE, a controlled benchmark in which models make deployment decisions with optional evidence that varies in retrieval cost, probability, severity, and presentation. Across GPT-5.5, o3, Claude Opus 4.8, and Claude Sonnet 4.6, we find distinct evidence-acquisition policies: Opus inspects nearly by default, o3 is the most skip-heavy and threshold-sensitive, and GPT-5.5 and Sonnet occupy intermediate regimes. Inspection increases strongly with severity and decreases with retrieval cost, whereas probability has much weaker behavioral influence: increasing the stated likelihood of a problem from 10% to 70% changes inspection by at most 21 percentage points. Despite these differences, Stage 1 rationales are dominated by expected-value reasoning across models. A cost-obligation decomposition further shows that avoidance is driven primarily by retrieval friction and explicit threats to the deployment payoff rather than by the remediation duties created by knowing. Counterfactual interventions reveal a further mismatch between behavior and explanation: evidence framing can strongly change decisions near the inspection boundary while going largely unmentioned, whereas probability is frequently cited despite having little causal influence. These results suggest that deployment-time safety depends not only on how models respond to known risks, but also on whether they acquire the evidence needed to know that acting is safe.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Dual-Frontier: When Can an Agent Trust Its World Model?

Dual-Frontier, a learning principle that admits a world-model-guided decision only when its predicted advantage exceeds a certified bound on decision-relevant world-model error; otherwise, evidence is allocated to world-model verification.

Hua-Tai Zhu, Qiang Chen, Zi-Qian Kou et al. · 1 citation
Review Open access Sep 2026

On the Cost of Conforming to Reviewers’ Expectations and the Potential Benefit of Prediction Competitions

Although academic peer review offers many important benefits, it can also impede scientific exploration. For instance, when reviewers share restrictive working assumptions, researchers may be incentivized to conform to them, even when alternative conjectures could better advance scientific understanding. We suggest mit...

Ido Erev, Adi Tarabeih, Rachel Barkan · 0 citations
Preprint Aug 2026

Bayesian Expected Uncertainty Reduction (B-EUR) Model: A Computational Account of What Makes Design Options Worth Trying

The B-EUR model provides a computational account of candidate-action evaluation within uncertainty-driven design activity and offers implications for constructing prototype sets, framing design problems, and organizing feedback to support informative exploration.

Shimon Honda, Takuma Miyaguchi, Koji Koizumi et al. · 0 citations
Preprint Aug 2026

Effort without Evidence

Advice can determine not only how past evidence is interpreted but whether new evidence will be produced. I study a sequence of decision makers who observe public success or failure but not one another's effort, and who pass costless causal assessments to their successors. Because a sender values success while underwei...

G. Lukyanov · 1 citation
Preprint Aug 2026

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings

The first benchmark that directly compares CoT monitorability under explicit-influence and implicit-influence regimes is introduced, suggesting that monitorability estimates obtained in explicit-influence settings may over-estimate monitorability, and that monitorability can be further decreased by well-intentioned dep...

Agatha Duzan, Asa Cooper Stickland · 2 citations
#artificial intelligence Preprint Sep 2026

Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding

This work identifies a failure mode not addressed by a stronger model: when the context contains an assertion by a party with an incentive toward optimism - here the sales representative, a witness recorded in the CRM - the model treats the assertion as evidence and clears deals the company's own records deem unaccepta...

Rahul Balakavi · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.