Preprint
Aug 2026
Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders
It is found that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states, but instead track model-internal similarity structure.
Nikolai Bolik, Lennart Stöpler, Artur Andrzejak
· 0 citations