Skip to content

First Token Matters: Understanding Safety Collapse in Large Reasoning Models

Sep 2026 · 0 citations · 45 references
Computer Science

TL;DR

It is found that the refusal-related signal of LRMs drops sharply at the first generated token under harmful queries, which is associated with unsafe response generation, and this work proposes SafeToken, a lightweight inference-time intervention that injects a learned continuous safety anchor precisely at reasoning onset.

Abstract

Large Reasoning Models (LRMs) exhibit strong problem-solving abilities, yet their safety alignment often degrades when handling harmful queries. Existing approaches to improving safety largely rely on additional training or preference optimization, while offering limited understanding of the internal mechanisms behind safety failures. In this work, we investigate this failure through a token-level positional analysis of refusal dynamics and identify a localized vulnerability at the onset of reasoning, which we term Onset Refusal Collapse (ORC). We find that the refusal-related signal of LRMs drops sharply at the first generated token under harmful queries, which is associated with unsafe response generation. Motivated by this finding, we propose SafeToken, a lightweight inference-time intervention that injects a learned continuous safety anchor precisely at reasoning onset. Despite updating only a single token embedding, SafeToken effectively mitigates ORC, improves safety on harmful-query benchmarks, and largely preserves reasoning utility. These results suggest that safety failures in LRMs can arise from a transient breakdown at the critical transition from understanding to generation.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

DeShortcut-Align: Decoupling Spurious Shortcuts for Robust Safety Alignment in Large Reasoning Models

DeShortcut-Align is proposed, a shortcut-decoupling alignment framework that reduces dependence on superficial cues that significantly improves robustness against template-stripping bypass attacks, substantially reduces over-refusal, and better preserves general-purpose reasoning capabilities, thereby mitigating the al...

Qi-Rui Liu, Yi-Chen Sun, Yan Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models

This work introduces DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses reasoning traces and final answers to quantify their safety inconsistency, and proposes SARA (Safety-Aware Reasoning Alignment), an RL-based method that rewards both safety-aware reasoning and safe final answers.

Xiang-Yu Zhou, S. Z. Zade, Rafi Ibn Sultan et al. · 0 citations
Preprint Aug 2026

TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models

TRACE is introduced, an evidence-grounded safety evaluation benchmark that covers the entire LRM inference pipeline: prompts, reasoning traces, and final responses, and reveals that safety judgment for reasoning traces is substantially more challenging than for prompts or final responses, and that current models strugg...

Zhenyu Wu, Siyu Chen, Chang-Chun Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

On the Efficiency-Safety Dilemma in Large Reasoning Models

Large reasoning models (LRMs) incur high inference costs, often mitigated by efficiency techniques like quantization and pruning. However, the impact of these techniques on model adversarial robustness remains largely unexplored. This study provides the first comprehensive analysis of the interplay between efficiency,...

Yifei Yang, Zou-Ying Cao, Xing-Rui Wang et al. · 0 citations
#natural language process... Preprint Sep 2026

When Confidence Rises Too Early: Detecting Shortcut Reasoning via Premature Answer Commitment

The reasoning trajectory of a Large Language Model (LLM) is often treated as a verbalized description of its internal reasoning. However, such trajectories can be unfaithful: a model may rely on shortcuts to reach an answer and then post-rationalize the decision with a seemingly coherent chain of thought. Detecting thi...

Zhao-Han Zhang, Jun-Jie Liu, Chengzhengxu Li et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.