ReliableNet is the only method certified within the JCW budget for every dataset and seed in distribution, when compared against baselines spanning ERM, post-hoc calibration, conformal risk control, and selective prediction, and selective prediction.
Abstract
A prediction that is both confident and wrong is a critical reliability failure because it can bypass abstention and human review precisely when the model is mistaken. Empirical risk minimization (ERM) controls average loss but not this failure directly, while calibration, uncertainty estimation, conformal risk control, and selective prediction methods target related reliability properties rather than bounding the joint failure event during training. We propose ReliableNet, which constrains the Joint Confident-Wrong (JCW) probability, the probability that a prediction is simultaneously confident and incorrect, below a user-specified risk budget $\alpha\in(0,1)$. We formulate this as a chance-constrained ERM problem, use a conservative smooth inner approximation whose population feasibility implies the original JCW constraint. Across four tabular and two image datasets, ReliableNet is the only method certified within the JCW budget for every dataset and seed in distribution, when compared against baselines spanning ERM, post-hoc calibration, conformal risk control, and selective prediction. Under demographic, ambiguity, spurious-correlation, novel-class, and covariate shifts, it achieves the lowest empirical JCW among the compared methods while remaining very competitive in accuracy, coverage, calibration, and selective prediction. Risk-coverage results further indicate that ReliableNet achieves better selective ranking than the benchmark methods on most datasets. Overall, ReliableNet provides a principled approach to trustworthy classification.
This thesis develops methods for improving calibration under label noise and studies calibration in unsupervised domain adaptation, where a model trained on a labeled source domain is adapted to an unlabeled target domain.
This work develops OCP with queries (OCPQ) by adapting the label efficient forecaster of Cesa-Bianchi, Lugosi, and Stoltz (2004) to the authors' setting, and develops OCP with queries (OCPQ) with queries in a way that encourages the learner to output small prediction sets while ensuring that the correct label is covere...
J. Skalse, Edoardo Pona, Osvaldo Simeone et al.· 0 citations
Reliable uncertainty quantification is essential for deploying deep learning models in high-stakes settings, where out-of-distribution and adversarial inputs can induce confident but unreliable predictions. Evidential Deep Learning provides efficient uncertainty estimates in a single forward pass, but can still assign...
Charmaine Barker, Daniel Bethell, Simos Gerasimou· 0 citations
RiskBlend is proposed, a classifier-agnostic prioritization framework that combines four complementary risk signals: historical failure patterns, prediction shift, decision-boundary shift, and neighborhood change that achieves the highest average APFD in all 80 dataset-classifier-scenario combinations.
Vibration-based bearing fault diagnosis informs maintenance decisions, but a predicted fault label is useful in practice only when the uncertainty of that prediction is quantified reliably. Existing deep-learning diagnostic models are often overconfident, poorly calibrated, and dependent on large labeled datasets, wh...
Jun-Chi Xu, Lu Liu, Guodong You et al.· Measurement science and tech...· 0 citations
Bounded predictive influence and reliability-guided geometry as complementary mechanisms for imbalanced learning with uncertain labels are supported as complementary mechanisms for imbalanced learning with uncertain labels.
M. Akhtar, J. Akarsh, M. Tanveer et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.