Skip to content
Preprint

Unmasking Removal-Budget Confounding: A Matched Operating-Point Evaluation Framework for Adaptive Data Cleaning

Aug 2026 · 0 citations · 18 references
Computer Science

TL;DR

Experiments on CIFAR-10 and ImageNet-100 demonstrate that most performance differences observed in naive evaluation shrink or vanish at low-to-moderate corruption when operating points are matched, highlighting that adaptive cleaning methods must be benchmarked at matched operating points to ensure performance gains reflect genuine corruption discrimination.

Abstract

Adaptive data-cleaning methods replace manual filtering thresholds with data-driven partitions. However, changing the partition granularity, the number of groups used to segment samples by estimated corruption risk, can implicitly shift the decision boundary and alter the overall number of removed samples. This creates a bias known as removal-budget confounding, where apparent gains in metrics like precision or false-positive rate reflect a smaller removal budget rather than superior corruption discrimination. To address this evaluation bias, we introduce an operating-point-aware evaluation framework that evaluates methods using matched-budget and matched-recall controls alongside threshold-independent metrics (AUROC and AUPRC). We test this framework on a multi-cue adaptive cleaner redesign featuring a reweighted learning-difficulty cue, an auxiliary Euclidean-distance cue, and increased partition granularity intended to isolate clean-but-difficult samples. While naive evaluations (assessing configurations at their own induced operating points) suggest substantial performance improvements for the redesign, these gains disappear once operating points are equalized. False-positive decomposition reveals that clean-but-difficult samples primarily drive error counts at low corruption rates, become threshold-dependent at moderate corruption, and contribute negligibly under severe corruption. Experiments on CIFAR-10 and ImageNet-100 demonstrate that most performance differences observed in naive evaluation shrink or vanish at low-to-moderate corruption when operating points are matched. True ranking advantages only remain in specific low-prevalence settings and in high-recall regions under severe corruption. These findings highlight that adaptive cleaning methods must be benchmarked at matched operating points to ensure performance gains reflect genuine corruption discrimination.

View source

Similar papers

#machine learning Preprint Sep 2026

A Free Knob: Decoupling Calibration and Predictive Skill in Threshold-Based Evaluation

A post-hoc transform fitted on held-out data and applied separately to each system is repurpose classical monotone calibration as a symmetric audit: a post-hoc transform fitted on held-out data and applied separately to each system.

Md Tanveer Hossain Munim, Bijoy Ahmed Saiem, Al-Amin Sany et al. · 0 citations

Principled Bias Detection and Mitigation across the ML Data Lifecycle

This dissertation treats data bias as a first-class data management problem and develops principled frameworks for detecting and correcting it across the ML data lifecycle, introducing Uniform Bias (UB), an interpretable intersectional measure with formal guarantees, and an ILP-based mitigation framework that models co...

Unknown authors · 0 citations
Book Open access Jun 2024

A Data-Centric Decomposition of Estimator Performance in Continuous Treatment Effect Estimation

This work analyzes current benchmarking practices and introduces a novel decomposition framework that disentangles the contribution of distinct data-generating components, such as confounding, dose distribution non-uniformity, and response surface complexity, to estimator performance.

Christopher Bockel-Rickermann, Daan Caljon, Toon Vanderschueren et al. · 2 citations
Preprint Aug 2026

Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection

It is shown the natural way to do this does not work, specify one that survives measurement, then finds that the correction making it work carries more variance than the null it is tested against, and that the correction making it work carries more variance than the null it is tested against.

Floriane C. M. Braun · 0 citations
Preprint Aug 2026

Hard Cases, Bad Labels: Testing Error Exposure and Error Location in Uncertainty Sampling Under Bounded Label Noise

Active learning can reduce labeling cost by selecting informative examples, but the most uncertain examples may also be the hardest to label correctly. This study tests whether uncertainty sampling fails because it acquires more corrupted labels or because errors concentrated in difficult regions are especially harmful...

John Myron Uy · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.