Skip to content

Local Sparsity Enables Unsupervised LLM Safety Detection

Sep 2026 · 0 citations · 68 references
Computer Science

TL;DR

This work proposes a framework for locally masked SAE-based anomaly detection, supported by theoretical justifications, and demonstrates their ability to capture meaningful safety information while using only 1-2% of SAE neurons for computation.

Abstract

Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and harm categories regularly arise, not captured by models trained in such a supervised fashion. An alternative approach is to view this problem through the lens of anomaly detection, namely, to rely solely on modeling safe data and flagging out-of-distribution inputs. However, LLM activations lie in a high-dimensional space, raising concerns about whether anomaly detection is statistically feasible. We show that, under the linear representation hypothesis (LRH), there may indeed be hope. In the LRH concept space, which is typically recovered via a sparse autoencoder (SAE), nearby points share a small common active support. Using this local sparsity insight, we propose a framework for locally masked SAE-based anomaly detection, supported by theoretical justifications. We validate it on various architectures and datasets, including both capability-testing datasets and safety-specific datasets. Finally, when we allow algorithms to use 1% out-of-distribution data for calibration, locally sparse methods achieve near-optimal performance, demonstrating their ability to capture meaningful safety information while using only 1-2% of SAE neurons for computation.

View source

Similar papers

Preprint Aug 2026

LLM as Detector: An In-context Learning Approach for Tabular Anomaly Detection

LLM-Detector is proposed, a framework that utilizes the in-context learning capacity of LLMs for structured, prompt-conditioned scoring synthesis, enabling LLMs to derive anomaly detection logic from structured normal-state knowledge.

Tu Nguyen, Dang Nguyen, T. D. Le et al. · 0 citations
2026

Sparsity-Controllable Normality Learning With Vision–Language Models for Scenario-Related Video Anomaly Detection

Video anomaly detection (VAD) is critical for automation systems and security surveillance. Recently, multimodal vision–language models (MLLMs) have attracted increasing attention due to their rich pre-trained knowledge and strong explainability. However, existing MLLM-based approaches struggle to adapt to real-world s...

Jiangyun Chen, Yuanjie Dang, Peng Chen et al. · 0 citations
Conference Open access Sep 2026

BERM: Low-Overhead Prompt-Injection Detection via In-Situ Benign Representation Modeling

BERM is introduced, a lightweight framework that performs in-situ detection by modeling a host LLM’s internal representations extracted during prefill, adding negligible overhead and reducing incremental inference overhead to near-zero.

Maihao Guo, Chaoyang Zhao, Jin-Qiao Wang · 0 citations
#machine learning Preprint Aug 2026

TrainSDC: Characterizing and Mitigating Silent Data Corruption in Large Language Model Training

This work presents the first systematic characterization of SDC vulnerability across major computation interfaces in both the forward and backward passes of Transformer training, and proposes TrainSDC, a characterization-guided protection framework consisting of Q/K-path recomputation, residual-gain monitoring, and exp...

Zhijie Xia, Haotian Xu, Si-Yu Yun et al. · 0 citations
#machine learning Preprint Sep 2026

Post-Anomaly Detection Inference for Deep SVDD

Deep Support Vector Data Description (Deep SVDD) has become a prominent framework for unsupervised anomaly detection by learning latent representations that compactly characterize normal data around a center. Despite its empirical success, anomaly decisions produced by Deep SVDD are typically made solely based on anoma...

Cao Le Cong Thanh, Vinh Quang Dang, Vo Nguyen Le Duy · 0 citations
#artificial intelligence Preprint Sep 2026

Faithful Dual-constrained Erasure for Robust LLM Safety Alignment

Machine unlearning has emerged as a crucial mechanism for removing hazardous knowledge and enforcing safety alignment in Large Language Models (LLMs). However, recent studies reveal a persistent security risk: unlearned models remain highly vulnerable to retraining attacks, where suppressed malicious behaviors rapidly...

Jia-Qing Li, Shi-De Zhou, Zhi-Bo Zhang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.