Experiments on three LRMs show that SaLT-DPO consistently reduces unsafe rates for both reasoning and answer segments while mitigating degradation in benign compliance and preserving general reasoning performance.
Abstract
Large Reasoning Models (LRMs) pose a dual-surface safety challenge: both intermediate reasoning traces and final answers can contain harmful content. Existing alignment methods often operate at the whole-response level, allowing unsafe reasoning to be masked by a safe-looking final answer. We propose Segment-aware Listwise Target DPO (SaLT-DPO), which addresses this gap through three mechanisms: (1) segment-aware listwise alignment that decomposes responses into reasoning and answer segments, independently scores each segment's safety, and aligns length-normalized segment rewards with soft target distributions over multiple candidates; (2) joint safety coherence regularization that applies a weakest-link principle to promote safety consistency across both segments; and (3) utility anchoring on benign prompts to mitigate over-refusal and reasoning degradation. Experiments on three LRMs show that SaLT-DPO consistently reduces unsafe rates for both reasoning and answer segments while mitigating degradation in benign compliance and preserving general reasoning performance. Ablation studies demonstrate the complementary contributions of its components.
It is found that the refusal-related signal of LRMs drops sharply at the first generated token under harmful queries, which is associated with unsafe response generation, and this work proposes SafeToken, a lightweight inference-time intervention that injects a learned continuous safety anchor precisely at reasoning on...
Yi-Zheng Yang, Hai-Ning Yu, Yue-Chen Wang et al.· 0 citations
TRACE is introduced, an evidence-grounded safety evaluation benchmark that covers the entire LRM inference pipeline: prompts, reasoning traces, and final responses, and reveals that safety judgment for reasoning traces is substantially more challenging than for prompts or final responses, and that current models strugg...
Zhenyu Wu, Siyu Chen, Chang-Chun Yang et al.· 0 citations
DeShortcut-Align is proposed, a shortcut-decoupling alignment framework that reduces dependence on superficial cues that significantly improves robustness against template-stripping bypass attacks, substantially reduces over-refusal, and better preserves general-purpose reasoning capabilities, thereby mitigating the al...
Qi-Rui Liu, Yi-Chen Sun, Yan Wang et al.· 0 citations
Table Question Answering (TableQA) requires reasoning over natural language questions and structured tables, and remains challenging due to noisy evidence and complex multi-step reasoning. Recent Large Language Model (LLM)-based approaches typically adopt decomposition–reasoning–validation pipelines that combine Chain-...
Zhen Yang, Zi-Wei Du, Ming-Han Zhang et al.· ACM Transactions on Informat...· 0 citations
The proposed TRACE-QA, a training-free multi-agent protocol that routes each instance to a sparse set of reasoning operators, constructs option-blind necessity constraints, audits every candidate in a structured elimination ledger, revisits risky eliminations through global risk-aware rescue, and aggregates role-specia...
Jia-Xin Lu, Hao Chen, Yan-Cheng Zhu et al.· Journal of King Saud Univers...· 0 citations
Large language models (LLMs) show strong reasoning ability, but their explanations can remain inconsistent, weakly grounded, or difficult to verify. We propose a verifier-guided explainable reasoning framework for transparent educational question answering that combines gold-anchored QLoRA, task-aware symbolic routing,...
Thi Kim Anh Vo, Nam-Tien Le, Thi Kim Anh Vo et al.· 0 citations