This work proposes a framework for locally masked SAE-based anomaly detection, supported by theoretical justifications, and demonstrates their ability to capture meaningful safety information while using only 1-2% of SAE neurons for computation.
Abstract
Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and harm categories regularly arise, not captured by models trained in such a supervised fashion. An alternative approach is to view this problem through the lens of anomaly detection, namely, to rely solely on modeling safe data and flagging out-of-distribution inputs. However, LLM activations lie in a high-dimensional space, raising concerns about whether anomaly detection is statistically feasible. We show that, under the linear representation hypothesis (LRH), there may indeed be hope. In the LRH concept space, which is typically recovered via a sparse autoencoder (SAE), nearby points share a small common active support. Using this local sparsity insight, we propose a framework for locally masked SAE-based anomaly detection, supported by theoretical justifications. We validate it on various architectures and datasets, including both capability-testing datasets and safety-specific datasets. Finally, when we allow algorithms to use 1% out-of-distribution data for calibration, locally sparse methods achieve near-optimal performance, demonstrating their ability to capture meaningful safety information while using only 1-2% of SAE neurons for computation.
LLM-Detector is proposed, a framework that utilizes the in-context learning capacity of LLMs for structured, prompt-conditioned scoring synthesis, enabling LLMs to derive anomaly detection logic from structured normal-state knowledge.
Tu Nguyen, Dang Nguyen, T. D. Le et al.· 0 citations
Video anomaly detection (VAD) is critical for automation systems and security surveillance. Recently, multimodal vision–language models (MLLMs) have attracted increasing attention due to their rich pre-trained knowledge and strong explainability. However, existing MLLM-based approaches struggle to adapt to real-world s...
Jiangyun Chen, Yuanjie Dang, Peng Chen et al.· IEEE Transactions on Informa...· 0 citations
BERM is introduced, a lightweight framework that performs in-situ detection by modeling a host LLM’s internal representations extracted during prefill, adding negligible overhead and reducing incremental inference overhead to near-zero.
Maihao Guo, Chaoyang Zhao, Jin-Qiao Wang· Proceedings of the Thirty-Fi...· 0 citations
This work presents the first systematic characterization of SDC vulnerability across major computation interfaces in both the forward and backward passes of Transformer training, and proposes TrainSDC, a characterization-guided protection framework consisting of Q/K-path recomputation, residual-gain monitoring, and exp...
Zhijie Xia, Haotian Xu, Si-Yu Yun et al.· 0 citations
Deep Support Vector Data Description (Deep SVDD) has become a prominent framework for unsupervised anomaly detection by learning latent representations that compactly characterize normal data around a center. Despite its empirical success, anomaly decisions produced by Deep SVDD are typically made solely based on anoma...
Cao Le Cong Thanh, Vinh Quang Dang, Vo Nguyen Le Duy· 0 citations
Machine unlearning has emerged as a crucial mechanism for removing hazardous knowledge and enforcing safety alignment in Large Language Models (LLMs). However, recent studies reveal a persistent security risk: unlearned models remain highly vulnerable to retraining attacks, where suppressed malicious behaviors rapidly...
Jia-Qing Li, Shi-De Zhou, Zhi-Bo Zhang et al.· 0 citations
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.