This work introduces SCALPEL, a contrastive sparse autoencoder designed to learn more selective forget features and shows theoretically that contrastive training promotes target-selective features and that the selection score controls expected background knowledge perturbation.
Abstract
Machine unlearning aims to remove targeted information while preserving a model's other abilities. In realistic settings, such as privacy requests under the EU GDPR, the target may be narrow, for example information associated with a single person. Behavioral forgetting alone may be insufficient, motivating interventions directly on internal representations. However, standard mechanistic-interpretability extractors are poorly selective for such targets. We identify an energy bias in reconstruction-based extraction, which favors dominant background structure over low-energy target-specific components. We introduce SCALPEL, a contrastive sparse autoencoder designed to learn more selective forget features. We show theoretically that contrastive training promotes target-selective features and that our selection score controls expected background knowledge perturbation. We validate SCALPEL experimentally on TOFU across Qwen, Llama, and Gemma, where it substantially improves over NMF and standard SAE interventions and is competitive with Gradient Difference and RMU, bridging mechanistic interpretability and fine-grained unlearning.
Machine unlearning aims to remove specific knowledge from a trained large language model (LLM) without retraining from scratch. Existing methods modify model weights via gradient ascent and its advances. While effective on certain benchmarks, these weight-based approaches exhibit a sharp forget-utility trade-off, where...
Pingzhi Li, Jinhao Duan, Vaishnav Tadiparthi et al.· 0 citations
Large Language Models (LLMs) can memorize and reproduce sensitive, copyrighted, or otherwise undesirable training content, creating privacy, safety, and regulatory concerns. Machine unlearning offers a practical alternative to full retraining, but many existing methods apply broad or fixed parameter updates that can de...
Ravi Ranjan, O. Kotevska, Agoritsa Polyzou· 0 citations
Machine unlearning (MU) aims to remove the influence of selected data from trained models, offering an efficient alternative to full retraining. With the rise of increasingly stringent privacy regulations, including the right to be forgotten, machine learning models must incorporate mechanisms that ensure compliance wh...
ADU is presented, a fine-grained, training-based framework that shifts unlearning from token erasure to contextual attention-pathway decoupling, and achieves the strongest aggregate performance among evaluated baselines on the TOFU and WMDP benchmarks.
Xun-Lei Chen, Qirui Ye, Yuang Li et al.· 0 citations
The ability of vision-language models (VLMs) to associate visual identities with biographical information creates a need for selective unlearning of personally identifiable information (PII) while preserving permitted knowledge about the same individual. This setting is challenging because both sensitive and retained i...
Si-Qi Goh, Cap Dang Xuan Kiet, Tat-Jen Cham et al.· 0 citations
A previously undocumented confound in how unlearning is evaluated on BatchNorm-based architectures is identified: a single forward pass over retain data, an operation that modifies no weight, can deterministically rewrite the model's normalization state and reverse the apparent surface-metric forgetting.
Aaryaman Kalani, Murari Mandal, Dhruv Kumar et al.· 1 citation
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.