Skip to content

Beyond Safe Answers: Segment-Aware Listwise Alignment for Reasoning Safety in Large Reasoning Models

Sep 2026 · 0 citations · 47 references
Computer Science

TL;DR

Experiments on three LRMs show that SaLT-DPO consistently reduces unsafe rates for both reasoning and answer segments while mitigating degradation in benign compliance and preserving general reasoning performance.

Abstract

Large Reasoning Models (LRMs) pose a dual-surface safety challenge: both intermediate reasoning traces and final answers can contain harmful content. Existing alignment methods often operate at the whole-response level, allowing unsafe reasoning to be masked by a safe-looking final answer. We propose Segment-aware Listwise Target DPO (SaLT-DPO), which addresses this gap through three mechanisms: (1) segment-aware listwise alignment that decomposes responses into reasoning and answer segments, independently scores each segment's safety, and aligns length-normalized segment rewards with soft target distributions over multiple candidates; (2) joint safety coherence regularization that applies a weakest-link principle to promote safety consistency across both segments; and (3) utility anchoring on benign prompts to mitigate over-refusal and reasoning degradation. Experiments on three LRMs show that SaLT-DPO consistently reduces unsafe rates for both reasoning and answer segments while mitigating degradation in benign compliance and preserving general reasoning performance. Ablation studies demonstrate the complementary contributions of its components.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

First Token Matters: Understanding Safety Collapse in Large Reasoning Models

It is found that the refusal-related signal of LRMs drops sharply at the first generated token under harmful queries, which is associated with unsafe response generation, and this work proposes SafeToken, a lightweight inference-time intervention that injects a learned continuous safety anchor precisely at reasoning on...

Yi-Zheng Yang, Hai-Ning Yu, Yue-Chen Wang et al. · 0 citations
Preprint Aug 2026

TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models

TRACE is introduced, an evidence-grounded safety evaluation benchmark that covers the entire LRM inference pipeline: prompts, reasoning traces, and final responses, and reveals that safety judgment for reasoning traces is substantially more challenging than for prompts or final responses, and that current models strugg...

Zhenyu Wu, Siyu Chen, Chang-Chun Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DeShortcut-Align: Decoupling Spurious Shortcuts for Robust Safety Alignment in Large Reasoning Models

DeShortcut-Align is proposed, a shortcut-decoupling alignment framework that reduces dependence on superficial cues that significantly improves robustness against template-stripping bypass attacks, substantially reduces over-refusal, and better preserves general-purpose reasoning capabilities, thereby mitigating the al...

Qi-Rui Liu, Yi-Chen Sun, Yan Wang et al. · 0 citations
Sep 2026

Focus on Evidence: Relational-Structure Enhances LLM Effectiveness in TableQA

Table Question Answering (TableQA) requires reasoning over natural language questions and structured tables, and remains challenging due to noisy evidence and complex multi-step reasoning. Recent Large Language Model (LLM)-based approaches typically adopt decomposition–reasoning–validation pipelines that combine Chain-...

Zhen Yang, Zi-Wei Du, Ming-Han Zhang et al. · 0 citations
Open access Aug 2026

TRACE-QA: Task-routed constraint elimination for auditable multi-agent question answering

The proposed TRACE-QA, a training-free multi-agent protocol that routes each instance to a sparse set of reasoning operators, constructs option-blind necessity constraints, audits every candidate in a structured elimination ledger, revisits risky eliminations through global risk-aware rescue, and aggregates role-specia...

Jia-Xin Lu, Hao Chen, Yan-Cheng Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR

Large language models (LLMs) show strong reasoning ability, but their explanations can remain inconsistent, weakly grounded, or difficult to verify. We propose a verifier-guided explainable reasoning framework for transparent educational question answering that combines gold-anchored QLoRA, task-aware symbolic routing,...

Thi Kim Anh Vo, Nam-Tien Le, Thi Kim Anh Vo et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.