A post-hoc transform fitted on held-out data and applied separately to each system is repurpose classical monotone calibration as a symmetric audit: a post-hoc transform fitted on held-out data and applied separately to each system.
Abstract
Many dense-prediction benchmarks evaluate rare events by pooling prediction and target over spatial blocks, thresholding each, and scoring the contingency table. At a fixed rare operating point, the max-pooled Critical Success Index (CSI) confounds spatial discrimination with amplitude calibration: sharp observations promote many blocks above threshold, while attenuated predictions from squared-error regression leave the same blocks below it. We repurpose classical monotone calibration as a symmetric audit: a post-hoc transform fitted on held-out data and applied separately to each system. The transform cannot reverse pixel ordering, so any contrast it reproduces cannot establish improved spatial ranking. On SEVIR, two released checkpoints of one architecture differ by -29.5% in extreme-threshold CSI before the control and by +5.3% after it. Across 450 pairwise contrasts among 6 systems, the difference in pooled frequency-bias deviation is associated with how far the CSI contrast moves under the control (r = +0.796), and 51 contrasts reverse sign. At CasCast's published extreme-event operating point, the cascade-over-backbone CSI gap falls from 0.1601 to 0.0339, a 78.8% reduction; the remaining gap stays positive. The effect persists when the transform is fitted on a window before the test period, and calibration also reveals advantages hidden by a better-calibrated baseline. On geostationary infrared imagery the relative gain grows as events become rarer, crowd counting reproduces the bias-gain relationship under patch-sum pooling, and semantic segmentation, where frequency bias is already near one, shows little average change. The confound therefore requires both a fixed operating point and a training regime that leaves the output miscalibrated there. We recommend reporting pooled frequency bias and a symmetric held-out FreeKnob Audit alongside rare-event pool-and-threshold scores.
SoftMCC is a calibration-sensitive MCC-family selector with bounded stability and utility evidence, coupling an MCC-specific calibrated identity with a tie-aware, shared-pool selection protocol.
Trusted monitoring has a cheap, trusted model score a stronger untrusted model's actions, and a diverse ensemble of them beats a single stronger monitor at matched cost. They are built by minimising average pairwise correlation, and that paper's twelve monitors shared one base model, leaving open what supplies the dive...
Many machine-learning systems set a threshold at a quantile of a calibration set: conformal predictors that promise 90% coverage by drawing their cutoff at the calibration set's 90th percentile, abstention gates that decline to answer when a model's score falls below the calibration set's tenth percentile, safety filte...
Experiments on CIFAR-10 and ImageNet-100 demonstrate that most performance differences observed in naive evaluation shrink or vanish at low-to-moderate corruption when operating points are matched, highlighting that adaptive cleaning methods must be benchmarked at matched operating points to ensure performance gains re...
Wei-Hsiang Chen, Pin-Hsuan Yu, Chenglin Fang et al.· 0 citations
Ranks depend on the observations used for comparison. Reusing those observations can add association between forecast and outcome rank contrasts even when the evaluated forecast and outcome stay fixed. We characterize assignments that preserve association between specified population-rank contrasts, including forecast...
PruneShift, an evaluation framework that separates broad predictive fidelity, fidelity near selector outputs, and the quality of the selected pruning decision, is introduced, showing why predictive fit, decision reliability, and pruning method quality require separate evidence.
Hao Ye, Gao-Peng Zhang· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Martin Trust Center Managing Director Bill Aulet introduces Dear Dreamer, a free platform for middle and high school students who want to learn about entrepreneurship.
Microsoft Research Blog· microsoft.comSep 30, 2026
Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.