Preprint
Jul 2026
Sparse Autoencoders for Interpretable Out-of-Distribution Detection
A novel approach that leverages sparse autoencoders (SAEs) to learn interpretable features from these intermediate activations and proposes a new OOD score derived from the cosine similarity between the sparse feature activations of a test sample and the mean activations of ID classes.
Ayush Karmacharya, Luke Luschwitz, Lucia Romero et al.
· 0 citations