Neuralyzing the Trace: Selective Representation-Level Unlearning with Contrastive Sparse Autoencoders
This work introduces SCALPEL, a contrastive sparse autoencoder designed to learn more selective forget features and shows theoretically that contrastive training promotes target-selective features and that the selection score controls expected background knowledge perturbation.
Itai Zehavi, Fanny Jourdan, Ulrich Aivodji
· 0 citations