Skip to content
Book Open access

NCCDA: Neuron-wise Class-Conditional Distribution Alignment for Deep Neural Network Repair

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 115-126 · 0 citations · 56 references

TL;DR

A novel general neural network repair paradigm termed NCCDA (Neuron-wise Class-Conditional Distribution Alignment), which theoretically prove a generalization error bound under small-sample settings based on Rademacher complexity, providing formal guarantees.

Abstract

Neural network repair aims to correct prediction failures caused by multiple security threats—such as backdoor attacks, natural corruptions, and safety property violations—through limited adjustments to model parameters. However, most existing repair methods rely on single-sample, point-to-point correction strategies, overlooking the statistical regularities of the feature space. As a result, they are highly sensitive to the scale of faulty samples and struggle to simultaneously achieve repair generalization and original performance preservation under small-sample settings. To address these limitations, we propose a novel general neural network repair paradigm termed NCCDA (Neuron-wise Class-Conditional Distribution Alignment). The method is grounded in a key insight: prediction failures fundamentally arise from neuron-level internal representations deviating from the high-likelihood regions corresponding to their true classes. NCCDA constructs neuron-wise class-conditional distribution references and formulates the repair process as a joint optimization of distribution alignment and structure preservation. By guiding abnormal representations back to high-likelihood regions while anchoring the structure of normal samples, the method enables efficient and adaptive repair without explicit neuron localization. We theoretically prove a generalization error bound under small-sample settings based on Rademacher complexity, providing formal guarantees. Extensive experiments across 7 benchmark datasets and 38 models, covering three categories of repair tasks, demonstrate that NCCDA consistently outperforms existing methods in repair effectiveness, generalization repair capability (Gene), and original accuracy preservation.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Let the Neurons Die: Exploiting ReLU-Induced Model Degradation

Rectified linear unit (ReLU) networks can suffer from dying neurons, where units with persistently negative pre-activations produce zero outputs, blocking gradients through their activations. To exploit this failure mode, we present three training-time availability attacks based on data ordering and poisoning. We begin...

Ke-Xin Li, Wen-Jun Qiu, Joshua Abraham et al. · 0 citations
Preprint Aug 2026

NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution

A fine-tuning-stage defense that simultaneously hardens LLMs against both attack classes by redistributing safety signals across a broader set of neurons, and provides a formal guarantee that NeuronGuard strictly reduces the attack success rate (ASR) upper bound.

Anjun Gao, Yueyang Quan, Yu Xia et al. · 2 citations
Open access Aug 2026

PGA-LLM: A Probability-Guided Alignment Large Language Model Framework for Fault Diagnosis

PGA-LLM, a novel fault diagnosis framework for industrial equipment that leverages large language models via probability-guided alignment and a progressive three-stage training scheme, encompassing encoder pre-training, interface optimization, and low-rank adaptation of Qwen2.

Tao Wang, Yanqiang Di, Shao-Chong Feng et al. · 0 citations
Open access Aug 2026

Human behavior alignment to improve the robustness of deep neural networks

This work proposes BrainTrain, a framework to create more robust DNNs through human behavior alignment and shows its utility in the context of object recognition and proposes Similarity Driven Label Smoothing (SDLS), a regularization method that scales BrainTrain to applications where it is difficult or expensive to co...

Bharath Anand, Sarada Krithivasan · 0 citations
Preprint Aug 2026

A Self-Explainable Deep Architecture for Security Applications

XSec is introduced, a self-explainable deep architecture developed for security applications that produces deterministic explanations for a fixed trained model and input and substantially reduces explanation latency compared with approximation-based and perturbation-based post-hoc methods.

Ananth Shreekumar, Jyun-Jhu Syu, Muslum Ozgur Ozmen et al. · 0 citations
Open access 2026

HIFN-Transformer: Learnable Information-Theoretic Parameters for Interpretable Deep Classification

HIFN-T is presented, a framework extending the Variational Information Bottleneck through four jointly learnable per-layer parameters: information retention, entropy budget, magnitude scaling, and global information gates that generalizes standard VIB as a special case and characterize the role of the entropy budget as...

Mohammed Tawfik · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.