Skip to content
Preprint

No Unique Minimizer, No Problem: On the Consistency of Robust Neural Classifiers

Aug 2026 · 0 citations · 29 references
Computer Science Mathematics

TL;DR

A consistency theory for robust neural classifiers based on the S-divergence family that requires no identifiability assumption is developed and it is proved that empirical S-divergence minimizers converge to the population-optimal equivalence class under mild regularity conditions.

Abstract

Neural network classifiers trained by cross-entropy minimization are highly sensitive to label noise and adversarial contamination. While robust alternatives offer bounded influence and resistance to corruption, their statistical foundations in the deep learning setting are insufficient due to a fundamental difficulty: neural parameterizations are non-identifiable, so the population loss minimizer is an equivalence class of parameters, not a unique point. We develop a consistency theory for robust neural classifiers based on the S-divergence family that requires no identifiability assumption. Casting training as stochastic optimization over a non-identifiable parameter space, we prove that empirical S-divergence minimizers converge to the population-optimal equivalence class under mild regularity conditions, and verify these conditions for three architecture choices. We further establish that limit points of the robust training algorithm are stationary points of the empirical objective. Experiments on vision and language benchmark datasets confirm that S-divergence training maintains clean-data accuracy while exhibiting performance competitive with existing robust methods.

View source

Similar papers

Preprint Jul 2026

On the robustness of noisy solutions in non-convex neural networks

Using a finite energy message-passing algorithm, it is demonstrated numerically that thermal noise enables effective generalization in the regime of constraint densities where both recovering the teacher and finding a zero temperature solution are computationally hard.

Enrico M. Malatesta, A. Passalacqua, Riccardo Zecchina · 0 citations
Preprint Jul 2026

Adversarial LassoNet: Robust Feature Selection via Stability-Driven Sparse Learning

Adversarial LassoNet is proposed, a stability-driven sparse feature selection framework that integrates input-space adversarial perturbations with the hierarchical sparsity mechanism of LassoNet and an NTK-inspired spectral analysis to characterize how perturbation-driven training can reduce gradient concentration.

Zhenghao Huang, Peicheng Xu, Junbiao Pang et al. · 0 citations
Preprint Jul 2026

Robustness Meets Uncertainty: Evidential Adversarial Training for Robust Selective Classification

Evidential Adversarial Training (EV-AT), which models uncertainty through a Dirichlet distribution and combines an evidence-based loss promoting clean accuracy and reliable uncertainty with a robust evidence-alignment loss matching clean and adversarial predictions in log Dirichlet-parameter space, is proposed.

Nicolas Sournac, Ahmed Baha Ben Jmaa, B. Braeckeveldt · 0 citations
Preprint Aug 2026

Bagging Robustly Learns VC Classes with Linear Sample Complexity

It is proved that VC classes are adversarially robustly learnable with sample complexity linear in the VC dimension $d$, providing an exponential improvement over the previous upper bound of Montasser, Hanneke, and Srebro (2019).

Omar Montasser · 0 citations
Preprint Aug 2026

Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification

Extensive experiments conducted on UCI and KEEL benchmark datasets demonstrate the superiority of the proposed IF-dRVFL and IF-edRVFL models over existing SOTA fuzzy and non-fuzzy approaches.

M. Sajid, A. Quadir, A. Rahaman et al. · 0 citations
Aug 2026

MC-SNN: Multicenter Stochastic Neural Network for Adversarially Robust Learning.

A multicenter learning method that leverages the advantage of stochastic neural networks (SNNs) for feature uncertainty learning and induces multiple centers for each class of samples in latent space to fit data more delicately, named the multicenter SNN (MC-SNN).

Meng Hu, Ran Wang, Yanting Guo et al. · 0 citations