Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· pp. 115-126· 0 citations· 18 references
Abstract
Neural network repair aims to correct prediction failures caused by multiple security threats—such as backdoor attacks, natural corruptions, and safety property violations—through limited adjustments to model parameters. However, most existing repair methods rely on single-sample, point-to-point correction strategies, overlooking the statistical regularities of the feature space. As a result, they are highly sensitive to the scale of faulty samples and struggle to simultaneously achieve repair generalization and original performance preservation under small-sample settings. To address these limitations, we propose a novel general neural network repair paradigm termed NCCDA (Neuron-wise Class-Conditional Distribution Alignment). The method is grounded in a key insight: prediction failures fundamentally arise from neuron-level internal representations deviating from the high-likelihood regions corresponding to their true classes. NCCDA constructs neuron-wise class-conditional distribution references and formulates the repair process as a joint optimization of distribution alignment and structure preservation. By guiding abnormal representations back to high-likelihood regions while anchoring the structure of normal samples, the method enables efficient and adaptive repair without explicit neuron localization. We theoretically prove a generalization error bound under small-sample settings based on Rademacher complexity, providing formal guarantees. Extensive experiments across 7 benchmark datasets and 38 models, covering three categories of repair tasks, demonstrate that NCCDA consistently outperforms existing methods in repair effectiveness, generalization repair capability (Gene), and original accuracy preservation.
VeRe is proposed, a verification-guided repair framework that leverages linear relaxation to precisely and efficiently estimate the repair significance of neurons and synthesizes ideal intervals that provide sound guarantees for correct behaviors, thereby facilitating surgical and targeted adjustments of neuron paramet...
Jia-Nan Ma, Wei Chen, Pengfei Yang et al.· ACM Transactions on Software...· 0 citations
A Neural-Collapse-Inspired Prioritization (NCIP) framework that replaces absolute confidence with cross-checkpoint prediction variability in the terminal training regime, where model geometry becomes highly structured and achieves strong performance in early fault discovery compared with competitive baselines.
Chun-Yu Liu, Ming-Yuan Li, Yang Li et al.· arXiv.org· 0 citations
A fine-tuning-stage defense that simultaneously hardens LLMs against both attack classes by redistributing safety signals across a broader set of neurons, and provides a formal guarantee that NeuronGuard strictly reduces the attack success rate (ASR) upper bound.
Anjun Gao, Yueyang Quan, Yu Xia et al.· 2 citations
PGA-LLM, a novel fault diagnosis framework for industrial equipment that leverages large language models via probability-guided alignment and a progressive three-stage training scheme, encompassing encoder pre-training, interface optimization, and low-rank adaptation of Qwen2.
Tao Wang, Yanqiang Di, Shao-Chong Feng et al.· Technologies· 0 citations
This work proposes BrainTrain, a framework to create more robust DNNs through human behavior alignment and shows its utility in the context of object recognition and proposes Similarity Driven Label Smoothing (SDLS), a regularization method that scales BrainTrain to applications where it is difficult or expensive to co...
Bharath Anand, Sarada Krithivasan· Frontiers in Artificial Inte...· 0 citations
Testing three delivery mechanisms–supervised fine-tuning, a dual-head classification loss, and reinforcement learning with a dense reward derived from the normalised penalty finds that supervised approaches consistently regress below the zero-shot baseline under distribution shift, while GRPO succeeds and generalises a...
Muntasir Adnan, Manile Srun, Carlos C. N. Kuhn· Machine Learning and Knowled...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.