Local Sparsity Enables Unsupervised LLM Safety Detection
This work proposes a framework for locally masked SAE-based anomaly detection, supported by theoretical justifications, and demonstrates their ability to capture meaningful safety information while using only 1-2% of SAE neurons for computation.
Xin Chen, Gil Kur, A. Shevchenko et al.
· 0 citations