A Neural-Collapse-Inspired Prioritization (NCIP) framework that replaces absolute confidence with cross-checkpoint prediction variability in the terminal training regime, where model geometry becomes highly structured and achieves strong performance in early fault discovery compared with competitive baselines.
Abstract
With the widespread deployment of deep neural networks (DNNs) in safety-critical domains, reducing the cost of model validation under limited testing budgets has become increasingly important. Existing test case prioritization techniques often rely on single-checkpoint confidence signals derived from output probabilities. However, DNNs can be confidently wrong, and the confidence margin between the predicted and competing classes is frequently small, which weakens early fault discovery. To address this limitation, we propose a Neural-Collapse-Inspired Prioritization (NCIP) framework that replaces absolute confidence with cross-checkpoint prediction variability in the terminal training regime, where model geometry becomes highly structured. NCIP introduces two key components. First, it selects an NC-guided representative subset of training checkpoints using an equiangularity score of classifier weights, quantified as the standard deviation of pairwise cosine similarities among class weight vectors. Second, it prioritizes test inputs by their prediction variability across the selected checkpoints, surfacing boundary-adjacent and failure-prone samples that are unstable under checkpoint-induced decision boundary shifts. Extensive experiments across multiple datasets and architectures show that NCIP achieves strong performance in early fault discovery compared with competitive baselines, with 1.5%–16.6% RAUC-ALL gains and 4.9%–20.6% RAUC-500 gains under the same testing budget. NCIP further attains the best average performance across all dataset-model pairs.
While Deep Neural Networks (DNNs) have achieved remarkable progress in cutting-edge domains, their inherent brittleness has become a growing concern. To ensure the reliability and safety of DNN-enabled software, DNN testing has emerged as an indispensable practice. Within this context, test input prioritization is esse...
Hao-Ran Li, Shi-Hai Wang, Bin Liu et al.· 0 citations
A novel general neural network repair paradigm termed NCCDA (Neuron-wise Class-Conditional Distribution Alignment), which theoretically prove a generalization error bound under small-sample settings based on Rademacher complexity, providing formal guarantees.
Liming Bao, Yan Wang, Tao Sun· Proceedings of the 32nd ACM...· 0 citations
RiskBlend is proposed, a classifier-agnostic prioritization framework that combines four complementary risk signals: historical failure patterns, prediction shift, decision-boundary shift, and neighborhood change that achieves the highest average APFD in all 80 dataset-classifier-scenario combinations.
Single-cell foundation models are increasingly adopted for downstream applications such as cell-type prediction. However, these predictions are often utilized without assessing their reliability, or by relying on a simple cutoff applied to the maximum softmax probability (MSP) derived from the classifier. This raises a...
This paper systematically analyzes failures encountered during large-scale RL training on the Huawei Ascend platform, summarizes representative failure types, and identifies three model-side factors relevant to fault reproduction.
Yikai Wang, Chuansai Zhou, Yuhang Zhou et al.· 0 citations
Recurrent Neural Networks (RNNs) have become a core component of modern intelligent software due to their strong ability to model temporal dependencies. As RNNs are increasingly deployed in safety-critical domains, ensuring their reliability is crucial. However, most existing testing techniques are designed for feedfor...
Xin-Yu Gao, Shuo-Xiao Zhang, Ming-Hui Wei et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.