It is found that the refusal-related signal of LRMs drops sharply at the first generated token under harmful queries, which is associated with unsafe response generation, and this work proposes SafeToken, a lightweight inference-time intervention that injects a learned continuous safety anchor precisely at reasoning onset.
Abstract
Large Reasoning Models (LRMs) exhibit strong problem-solving abilities, yet their safety alignment often degrades when handling harmful queries. Existing approaches to improving safety largely rely on additional training or preference optimization, while offering limited understanding of the internal mechanisms behind safety failures. In this work, we investigate this failure through a token-level positional analysis of refusal dynamics and identify a localized vulnerability at the onset of reasoning, which we term Onset Refusal Collapse (ORC). We find that the refusal-related signal of LRMs drops sharply at the first generated token under harmful queries, which is associated with unsafe response generation. Motivated by this finding, we propose SafeToken, a lightweight inference-time intervention that injects a learned continuous safety anchor precisely at reasoning onset. Despite updating only a single token embedding, SafeToken effectively mitigates ORC, improves safety on harmful-query benchmarks, and largely preserves reasoning utility. These results suggest that safety failures in LRMs can arise from a transient breakdown at the critical transition from understanding to generation.
Experiments on three LRMs show that SaLT-DPO consistently reduces unsafe rates for both reasoning and answer segments while mitigating degradation in benign compliance and preserving general reasoning performance.
Jungmin Yun, Junehyoung Kwon, Hayeong Ryu et al.· 0 citations
DeShortcut-Align is proposed, a shortcut-decoupling alignment framework that reduces dependence on superficial cues that significantly improves robustness against template-stripping bypass attacks, substantially reduces over-refusal, and better preserves general-purpose reasoning capabilities, thereby mitigating the al...
Qi-Rui Liu, Yi-Chen Sun, Yan Wang et al.· 0 citations
This work introduces DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses reasoning traces and final answers to quantify their safety inconsistency, and proposes SARA (Safety-Aware Reasoning Alignment), an RL-based method that rewards both safety-aware reasoning and safe final answers.
Xiang-Yu Zhou, S. Z. Zade, Rafi Ibn Sultan et al.· 0 citations
TRACE is introduced, an evidence-grounded safety evaluation benchmark that covers the entire LRM inference pipeline: prompts, reasoning traces, and final responses, and reveals that safety judgment for reasoning traces is substantially more challenging than for prompts or final responses, and that current models strugg...
Zhenyu Wu, Siyu Chen, Chang-Chun Yang et al.· 0 citations
Large reasoning models (LRMs) incur high inference costs, often mitigated by efficiency techniques like quantization and pruning. However, the impact of these techniques on model adversarial robustness remains largely unexplored. This study provides the first comprehensive analysis of the interplay between efficiency,...
Yifei Yang, Zou-Ying Cao, Xing-Rui Wang et al.· 0 citations
The reasoning trajectory of a Large Language Model (LLM) is often treated as a verbalized description of its internal reasoning. However, such trajectories can be unfaithful: a model may rely on shortcuts to reach an answer and then post-rationalize the decision with a seemingly coherent chain of thought. Detecting thi...
Zhao-Han Zhang, Jun-Jie Liu, Chengzhengxu Li et al.· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.