Skip to content

The BatchNorm Illusion: Diagnosing Normalization Artifacts in Machine Unlearning Evaluation

Sep 2026 · 1 citation · 31 references
Computer Science

TL;DR

A previously undocumented confound in how unlearning is evaluated on BatchNorm-based architectures is identified: a single forward pass over retain data, an operation that modifies no weight, can deterministically rewrite the model's normalization state and reverse the apparent surface-metric forgetting.

Abstract

Approximate machine unlearning aims to remove the influence of specific training data from a trained model without retraining from scratch. We identify a previously undocumented confound in how unlearning is evaluated on BatchNorm-based architectures: a single forward pass over retain data, an operation that modifies no weight, can deterministically rewrite the model's normalization state and reverse the apparent surface-metric forgetting. We formalize this operation as a weight-preserving fixed-point operator and prove that any pre-versus-post gap it induces is provably attributable to BN running statistics rather than to any modification the unlearning method made to the weights. This attribution claim cleanly separates measurement failure (BN artifact) from encoder failure (residual weight-encoded information, recently documented in concurrent work), and the same operator framework yields a unique decomposition of linear-probe elevation into BN-measurement-bias and encoder-geometry components. Empirically, the artifact reverses headline forget accuracy by up to 78 pp across nine evaluated methods on standard benchmarks; an attacker with as few as 10 unlabeled images recovers most of the masked accuracy; and a strict GroupNorm control reduces the artifact to zero across all methods. The tested membership-inference attacks change little under recalibration, locating the observed evaluation failure in forget accuracy and linear probing.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Unmerge: Efficient Machine Unlearning via Task Arithmetic

Approximate machine unlearning seeks to remove the influence of a forget set from a trained model without full retraining. Existing gradient-based methods require data-dependent hyperparameter search, struggle when forget and retain knowledge are entangled, and offer little insight into where unlearning actually happen...

Hao-Ran Tang, Andrew Tan, Rajiv Khanna · 0 citations
Preprint Aug 2026

Cross-Domain Generalization in Machine Unlearning via Label-Conditioned Energy Magnitude Regularization

This paper studies what happens to the rest of the model when a class is forgotten, using a label-conditioned energy-based model (EBM) that assigns per-class energies, making the effect directly observable.

Syed Ali Ahmed, Syed Bilal Ahsan, Muhammad Zaigham Zaheer National University of Computer et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Do Influence-Derived Data Perturbations Enable Machine Unlearning? A Controlled Study of Three Plausible Roles

We evaluate Deep Perturbation Learning (DPL), which perturbs training images and labels along influence-derived directions, in three roles in which prior work has positioned it for machine unlearning: a direct deletion signal (the strongest claim), a utility-preserving regularizer, and a warm start for adversarial unle...

Chen Wu, Chrispine Kambimbi, Qi Zeng et al. · 0 citations
#machine learning Preprint Sep 2026

Test-Time Unlearning via Sparse Autoencoder

Machine unlearning aims to remove specific knowledge from a trained large language model (LLM) without retraining from scratch. Existing methods modify model weights via gradient ascent and its advances. While effective on certain benchmarks, these weight-based approaches exhibit a sharp forget-utility trade-off, where...

Pingzhi Li, Jinhao Duan, Vaishnav Tadiparthi et al. · 0 citations
2026

DeepU: Deeper Granular Within-Layer Machine Unlearning

Machine unlearning (MU) aims to remove the influence of selected data from trained models, offering an efficient alternative to full retraining. With the rise of increasingly stringent privacy regulations, including the right to be forgotten, machine learning models must incorporate mechanisms that ensure compliance wh...

Anudeep Vurity, Zhi-Sheng Yan, Massimiliano Albanese · 0 citations
#artificial intelligence Preprint Sep 2026

Neuralyzing the Trace: Selective Representation-Level Unlearning with Contrastive Sparse Autoencoders

This work introduces SCALPEL, a contrastive sparse autoencoder designed to learn more selective forget features and shows theoretically that contrastive training promotes target-selective features and that the selection score controls expected background knowledge perturbation.

Itai Zehavi, Fanny Jourdan, Ulrich Aivodji · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.