Emotional expression can influence the safety decisions of large language models (LLMs), offering a potential avenue for improving safety alignment. Existing studies have mainly focused on how emotional expressions facilitate attacks under harmful requests, while overlooking their effects on benign requests. We find th...
Shu-Yi Miao, Yao-Jin Ma, Chen-Hang Cui et al.· 0 citations
Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence. Currently, mechanistic interpretability studies on multilingual safety are largely confined to local components, such as isolated neurons. However, this st...
Shu-Yi Miao, Wang-Jie Qiu, Pengyang Shao et al.· 0 citations
This work proposes aligning cross-lingual thoughts and responses (ACTR), a framework that improves multilingual safety alignment by strengthening the use of existing safety reasoning, and devise neuron-selective consistency optimization (NSCO), which uses a frozen judge model to reward agreement between the safety cate...
Xian-Hui Zhang, Jian Yu, Cheng-Yu Xie et al.· 0 citations
A neuron-level cross-dimensional safety alignment framework driven by modality- and language-shared safety neurons (MLS-Neurons) that significantly outperforms state-of-the-art approaches across diverse multilingual and multimodal safety benchmarks while preserving general utility.
Enyi Shi, Fei Shen, Chuancheng Shi et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.