This work introduces a retain-aware localization method that considers parameter importance to both forgotten and retained data, and introduces a retain-similar evaluation set, constructed using cosine similarity in the model embedding space, to directly measure collateral damage.
Abstract
Machine unlearning removes the influence of user-specified training examples from a trained model, avoiding the need to retrain it from scratch. Localization-based methods improve unlearning efficiency by identifying a subset of influential model parameters. However, existing approaches select parameters based solely on forget-set importance, neglecting their role in retained dataset and often causing collateral damage to semantically similar retained examples. We address this limitation with a retain-aware localization method that considers parameter importance to both forgotten and retained data. We also introduce a retain-similar evaluation set, constructed using cosine similarity in the model embedding space, to directly measure collateral damage. Across eleven experimental settings on CIFAR-10 dataset and ResNet18 model, our method consistently reduces collateral damage while improving standard unlearning metrics, demonstrating the effectiveness of retain-aware localization for similarity-aware machine unlearning.
Driven by stringent privacy regulations and data deletion requirements, machine unlearning has emerged as a critical field focused on selectively removing the influence of specific data from the pre-trained model. In this paper, we focus on achieving the unlearning objective while maintaining better model accuracy. Fir...
GRACE, a gradient-guided coreset selection method that constructs both forget and retain sets for LLM unlearning, improves model utility while maintaining comparable forget quality, with particularly consistent gains over prior gradient-based selection methods.
Praveen Bushipaka, Andrea D'Angelo, Lucia C. Passaro et al.· 0 citations
Machine unlearning (MU) aims to remove the influence of selected data from trained models, offering an efficient alternative to full retraining. With the rise of increasingly stringent privacy regulations, including the right to be forgotten, machine learning models must incorporate mechanisms that ensure compliance wh...
Inspired by Muon, this work adopts the spectral view for unlearning and proposes Spectral Saliency Unlearning (SSU), a thresholding approach that thresholds weak singular components and updates only those directions supported by a confident unlearning signal.
Cedar Site Bai, Amber Yijia Zheng, Raymond A. Yeh et al.· 0 citations
This paper studies what happens to the rest of the model when a class is forgotten, using a label-conditioned energy-based model (EBM) that assigns per-class energies, making the effect directly observable.
Syed Ali Ahmed, Syed Bilal Ahsan, Muhammad Zaigham Zaheer National University of Computer et al.· 0 citations
Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this issue, existing unlearning methods typically rely on training-based parameter updates, such as gradient ascent and its variants, to delete targeted content while preservin...
Pu-Ning Yang, Qi-Zhou Wang, Jun-Chi Yu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.