A novel calibration method is proposed, SimCal, which uses synthetic data generated from the model development data in conjunction with marginal statistics from the calibration cohort to address the challenge of calibrating statistical prediction models for binary outcomes when training data is lacking.
Abstract
Statistical prediction models for binary outcomes are becoming increasingly popular. One significant challenge is calibrating these models to suit the characteristics of a target population that is structurally different from the original population. Calibration is especially challenging when there is no training data available from the target population. To address this problem, we propose a novel calibration method, SimCal, which uses synthetic data generated from the model development data in conjunction with marginal statistics from the calibration cohort. We show that expert judgment modeling (EJM) may be used for calibration if cross-sectional data from the target population are available comprising expert judgments about the potential outcome and the covariates. We describe three alternative calibration approaches when calibration data are lacking: similarity-binning averaging (SBA), adaptive calibration of predictions (ACP), and Elkan calibration. In a simulation study, we compare SBA, ACP, Elkan calibration, and SimCal. R code for applying these methods is provided from the re-analysis of data on coronary artery disease. We illustrate all 5 calibration approaches with a real data set for predicting functional outcome after stroke and all approaches but EJM in the re-analysis of the Cleveland Clinic data. None of the approaches performed convincingly well in all situations. SimCal performed well when model parameters were correctly specified. EJM failed on the stroke data. Further research is urgently required for calibration in the absence of calibration data.
Assessing the goodness-of-fit of a logistic regression model is a critical prerequisite before the model is used for inference. However, goodness-of-fit (GOF) tests such as the chi-square and deviance tests often give invalid results when the data are"sparse"-- a common issue with continuous predictors like age or weig...
Martingale posteriors and related predictive resampling methods replace the likelihood--prior pair used within Bayesian inference with a predictive model for future observations. These methods are simple to implement and increasingly popular due to their computational efficiency, but little is known about their ability...
Hui Wang, Edwin Fong, David T. Frazier· 0 citations
Cross-fitted estimators that transport information from the two labeled sources through source-specific density ratios are proposed that establish asymptotically linear inference for TPR and FPR, consistency and pointwise inference for the ROC curve, and asymptotically normal inference for AUC.
This work analyzes current benchmarking practices and introduces a novel decomposition framework that disentangles the contribution of distinct data-generating components, such as confounding, dose distribution non-uniformity, and response surface complexity, to estimator performance.
Christopher Bockel-Rickermann, Daan Caljon, Toon Vanderschueren et al.· Proceedings of the 32nd ACM...· 2 citations
Prediction models can perform poorly when the deployment population differs from the training population. Data from the target population would help, but individual-level target data may be inaccessible because of access restrictions or reporting conventions. We consider a multi-resolution setting in which individual-l...
The findings corroborate the concern that standard resampling methods often yield biased GE estimates in nonstandard settings, underscoring the importance of tailored GE estimation.
R. Hornung, Malte Nalenz, Lennart Schneider et al.· Statistical Science· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.