Skip to content

Predicting Privacy Leakage from Weight Spectral Density

Sep 2026 · 0 citations · 27 references
Computer Science

TL;DR

The results indicate that neural network spectra may contain information about privacy leakage that is not fully captured by conventional measures of overfitting, motivating spectral analysis as a promising direction for scalable privacy auditing.

Abstract

Membership inference attacks (MIAs) are widely used to audit the privacy disclosure risk of machine learning models, however current state-of-the-art attacks require training computationally expensive shadow models, making large-scale privacy evaluation impractical. In this work, we investigate whether inexpensive spectral metrics derived from the heavy-tailed self-regularisation framework can serve as proxies for MIA vulnerability. We evaluate several WeightWatcher spectral metrics on image and tabular classification tasks and compare their relationship with MIA privacy leakage against conventional measures of generalisation. Across datasets, stable rank exhibits a strong positive correlation with overall MIA success, while Log alpha-Norm shows a consistent negative correlation with MIA vulnerability at the low false-positive regime. These associations are observed to be stronger than those obtained using the generalisation gap. The results indicate that neural network spectra may contain information about privacy leakage that is not fully captured by conventional measures of overfitting, motivating spectral analysis as a promising direction for scalable privacy auditing.

View source

Similar papers

Book Open access Aug 2026

Efficient Privacy Auditing for Generative Model via Local Information

This work proposes an empirical privacy assessment framework that leverages local image information, instead of treating images as indivisible wholes, to improve the distinguishability of privacy leakage signals and further leverage the image inpainting interface of diffusion models to perform region-focused auditing i...

Jingnan Xu, Leixia Wang, Xiaofeng Meng · 0 citations
#machine learning Preprint Sep 2026

Membership Inference via Pairwise Likelihood Ratios

Membership inference attacks (MIAs) are the standard tool for auditing the privacy risks of machine learning models. Given a query point, an MIA aims to determine whether that point was used to train the target model. In practice, such inference must rely on the statistical signals exposed by the model's outputs, such...

Sheng-Jie Niu, Ze-Bin Yun, Ye-Heng Ge et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Sharp Transition in Data Reconstruction under Differential Privacy

Data reconstruction attacks have empirically been successful in recovering training samples from learned models, raising privacy concerns and motivating defenses with guarantees that remain valid against future threats. While differential privacy (DP) provides formal protection, choosing the privacy budget remains a ch...

Max Cairney-Leeming, Simone Bombari, Marco Mondelli · 0 citations
Review Sep 2026

Privacy-preserving methodologies against privacy attacks on deep learning: a survey

The analysis finds that noise-based methods such as DP remain the practical baseline but offer only partial protection; cryptographic approaches provide stronger theoretical guarantees at substantially higher cost; and LLM/multimodal leakage remains an urgent, under-benchmarked gap.

Subhasish Ghosh, A. K. Mandal · 0 citations
Conference Open access Aug 2026

CutClean: Neural Network Pruning for Privacy-Preserving Inference

CutClean is a privacy-aware pruning method that allows to reduce privacy information flow through the network, while increasing its sparsity, and effectively minimizes private information flow while achieving high sparsity rates and preserving classification target accuracy.

Leonardo Magliolo, Vito Paolo Pastore, Giuseppe Valenzise et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.