Skip to content
Open access

Test Case Prioritization for DNNs via Neural Collapse Instability

Jul 2026 · arXiv.org · Vol 3, pp. 2113 - 2135 · 0 citations · 65 references
Computer Science

TL;DR

A Neural-Collapse-Inspired Prioritization (NCIP) framework that replaces absolute confidence with cross-checkpoint prediction variability in the terminal training regime, where model geometry becomes highly structured and achieves strong performance in early fault discovery compared with competitive baselines.

Abstract

With the widespread deployment of deep neural networks (DNNs) in safety-critical domains, reducing the cost of model validation under limited testing budgets has become increasingly important. Existing test case prioritization techniques often rely on single-checkpoint confidence signals derived from output probabilities. However, DNNs can be confidently wrong, and the confidence margin between the predicted and competing classes is frequently small, which weakens early fault discovery. To address this limitation, we propose a Neural-Collapse-Inspired Prioritization (NCIP) framework that replaces absolute confidence with cross-checkpoint prediction variability in the terminal training regime, where model geometry becomes highly structured. NCIP introduces two key components. First, it selects an NC-guided representative subset of training checkpoints using an equiangularity score of classifier weights, quantified as the standard deviation of pairwise cosine similarities among class weight vectors. Second, it prioritizes test inputs by their prediction variability across the selected checkpoints, surfacing boundary-adjacent and failure-prone samples that are unstable under checkpoint-induced decision boundary shifts. Extensive experiments across multiple datasets and architectures show that NCIP achieves strong performance in early fault discovery compared with competitive baselines, with 1.5%–16.6% RAUC-ALL gains and 4.9%–20.6% RAUC-500 gains under the same testing budget. NCIP further attains the best average performance across all dataset-model pairs.

Read PDF

Similar papers

Preprint Sep 2026

When Ambiguity Meets Atypicality: Dual-Perspective Test Input Prioritization for DNNs

While Deep Neural Networks (DNNs) have achieved remarkable progress in cutting-edge domains, their inherent brittleness has become a growing concern. To ensure the reliability and safety of DNN-enabled software, DNN testing has emerged as an indispensable practice. Within this context, test input prioritization is esse...

Hao-Ran Li, Shi-Hai Wang, Bin Liu et al. · 0 citations
Book Open access Aug 2026

NCCDA: Neuron-wise Class-Conditional Distribution Alignment for Deep Neural Network Repair

A novel general neural network repair paradigm termed NCCDA (Neuron-wise Class-Conditional Distribution Alignment), which theoretically prove a generalization error bound under small-sample settings based on Rademacher complexity, providing formal guarantees.

Liming Bao, Yan Wang, Tao Sun · 0 citations
#artificial intelligence Review Aug 2026

RiskBlend: A Multi-Signal Framework for Test Input Prioritization in Machine Learning Regression Testing

RiskBlend is proposed, a classifier-agnostic prioritization framework that combines four complementary risk signals: historical failure patterns, prediction shift, decision-boundary shift, and neighborhood change that achieves the highest average APFD in all 80 dataset-classifier-scenario combinations.

Madhusudan Srinivasan, Namith Nishal Raphae · 0 citations
Review Open access Sep 2026

Cross-domain confidence reliability and remappability of frozen single-cell representations

Single-cell foundation models are increasingly adopted for downstream applications such as cell-type prediction. However, these predictions are often utilized without assessing their reliability, or by relying on a simple cutoff applied to the maximum softmax probability (MSP) derived from the classifier. This raises a...

Guang-Zheng Weng, Dan-Fei Zhu, Yu-Fei Zhao et al. · 0 citations
Preprint Aug 2026

MoE Proxy Models for Low-Cost Failure Reproduction and Diagnosis in LLM RL Post-Training

This paper systematically analyzes failures encountered during large-scale RL training on the Huawei Ascend platform, summarizes representative failure types, and identifies three model-side factors relevant to fault reproduction.

Yikai Wang, Chuansai Zhou, Yuhang Zhou et al. · 0 citations
Open access Oct 2026

StateTree: A Tree-Based Modeling Approach for Fault Detection in Recurrent Neural Networks

Recurrent Neural Networks (RNNs) have become a core component of modern intelligent software due to their strong ability to model temporal dependencies. As RNNs are increasingly deployed in safety-critical domains, ensuring their reliability is crucial. However, most existing testing techniques are designed for feedfor...

Xin-Yu Gao, Shuo-Xiao Zhang, Ming-Hui Wei et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.