Experiments on CIFAR-10 and ImageNet-100 demonstrate that most performance differences observed in naive evaluation shrink or vanish at low-to-moderate corruption when operating points are matched, highlighting that adaptive cleaning methods must be benchmarked at matched operating points to ensure performance gains reflect genuine corruption discrimination.
Abstract
Adaptive data-cleaning methods replace manual filtering thresholds with data-driven partitions. However, changing the partition granularity, the number of groups used to segment samples by estimated corruption risk, can implicitly shift the decision boundary and alter the overall number of removed samples. This creates a bias known as removal-budget confounding, where apparent gains in metrics like precision or false-positive rate reflect a smaller removal budget rather than superior corruption discrimination. To address this evaluation bias, we introduce an operating-point-aware evaluation framework that evaluates methods using matched-budget and matched-recall controls alongside threshold-independent metrics (AUROC and AUPRC). We test this framework on a multi-cue adaptive cleaner redesign featuring a reweighted learning-difficulty cue, an auxiliary Euclidean-distance cue, and increased partition granularity intended to isolate clean-but-difficult samples. While naive evaluations (assessing configurations at their own induced operating points) suggest substantial performance improvements for the redesign, these gains disappear once operating points are equalized. False-positive decomposition reveals that clean-but-difficult samples primarily drive error counts at low corruption rates, become threshold-dependent at moderate corruption, and contribute negligibly under severe corruption. Experiments on CIFAR-10 and ImageNet-100 demonstrate that most performance differences observed in naive evaluation shrink or vanish at low-to-moderate corruption when operating points are matched. True ranking advantages only remain in specific low-prevalence settings and in high-recall regions under severe corruption. These findings highlight that adaptive cleaning methods must be benchmarked at matched operating points to ensure performance gains reflect genuine corruption discrimination.
A post-hoc transform fitted on held-out data and applied separately to each system is repurpose classical monotone calibration as a symmetric audit: a post-hoc transform fitted on held-out data and applied separately to each system.
Md Tanveer Hossain Munim, Bijoy Ahmed Saiem, Al-Amin Sany et al.· 0 citations
This dissertation treats data bias as a first-class data management problem and develops principled frameworks for detecting and correcting it across the ML data lifecycle, introducing Uniform Bias (UB), an interpretable intersectional measure with formal guarantees, and an ILP-based mitigation framework that models co...
This work analyzes current benchmarking practices and introduces a novel decomposition framework that disentangles the contribution of distinct data-generating components, such as confounding, dose distribution non-uniformity, and response surface complexity, to estimator performance.
Christopher Bockel-Rickermann, Daan Caljon, Toon Vanderschueren et al.· Proceedings of the 32nd ACM...· 2 citations
It is shown the natural way to do this does not work, specify one that survives measurement, then finds that the correction making it work carries more variance than the null it is tested against, and that the correction making it work carries more variance than the null it is tested against.
Under homoscedastic retrieval noise, it is shown that retrieval failure decays exponentially with task separation relative to noise, and explicit finite-sample conditions under which RAG-FT achieves lower risk than both target-only and full-corpus training are derived.
Active learning can reduce labeling cost by selecting informative examples, but the most uncertain examples may also be the hardest to label correctly. This study tests whether uncertainty sampling fails because it acquires more corrupted labels or because errors concentrated in difficult regions are especially harmful...
John Myron Uy· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.