Decoding One Safety Trigger Token for Balancing Safety and Usability in Large Language Models
D-STT, a simple yet effective defense algorithm that identifies and explicitly decodes safety trigger tokens of the given safety-aligned LLM to activate the model's learned safety patterns, effectively preserves model usability by introducing minimum intervention in the decoding process.
Hao-Ran Gu, Han-Ding Wang, Yi Mei et al.
· 0 citations