This survey provides a structured, critical review of methods for Uncertainty Quantification in deep learning, scoped to ensemble-based and approximate Bayesian approaches and the measures used to summarize their outputs.
Abstract
The deployment of deep neural networks in safety-critical domains demands reliable estimates of predictive confidence, yet conventional architectures lack principled uncertainty quantification. This survey provides a structured, critical review of methods for Uncertainty Quantification (UQ) in deep learning, scoped to ensemble-based and approximate Bayesian approaches and the measures used to summarize their outputs. Relative to existing UQ surveys, our contribution is depth on efficient ensemble approximations and single-pass methods, and a unified treatment that separates the method producing a predictive distribution from the measure that summarizes its uncertainty. We organize methods into five families: Bayesian neural networks, Monte Carlo Dropout, deep ensembles, efficient ensemble approximations, and last-layer or single-pass approaches. We situate adjacent work on evidential and prior networks, conformal prediction, and post-hoc calibration, together with the decision-time tasks of out-of-distribution detection and selective prediction. For each, we examine theoretical motivation, implementation, empirical performance, and limitations. We then review ensemble diversity theory and uncertainty measures and their decompositions, contrasting the entropy decomposition with pairwise divergence measures, and consolidate evaluation methodology so that our qualitative comparisons share a common basis. We close with a brief treatment of uncertainty in large language models and open research directions, including efficient epistemic measures for classification, last-layer diversity, diversity and calibration under shift, and hybrid architectures.
This thesis develops methods for improving calibration under label noise and studies calibration in unsupervised domain adaptation, where a model trained on a labeled source domain is adapted to an unlabeled target domain.
Experimental evaluations show that this method outperforms traditional deep learning methods in terms of uncertainty quantification accuracy, training stability, and resource utilization, providing a technical foundation for the development of deep learning models in high-reliability systems.
Zhanyi Wei· International Conference on...· 0 citations
This work presents an empirical analysis of UQ in deep learning models, focusing on genomics applications, and shows that Bayesian Neural Networks are better at capturing uncertainty caused by strong class imbalance and out-of-distribution data in genomics, despite their computational disadvantages.
Extensive experiments conducted on UCI and KEEL benchmark datasets demonstrate the superiority of the proposed IF-dRVFL and IF-edRVFL models over existing SOTA fuzzy and non-fuzzy approaches.
M. Sajid, A. Quadir, A. Rahaman et al.· 0 citations
Con-formal prediction can be used to enrich existing techniques, with post-hoc uncertainty quantification in a plug and play fashion, and consistently outperforms the state of the art across six out of eight real world scenarios.
S. Jahns, Johannes Dorn, Max Weber et al.· 0 citations
ReliableNet is the only method certified within the JCW budget for every dataset and seed in distribution, when compared against baselines spanning ERM, post-hoc calibration, conformal risk control, and selective prediction, and selective prediction.