Skip to content
Preprint

Benchmarking Goodness-of-Fit and Calibration Algorithms for Logistic Regression Classifiers: A Large-Scale Simulation Study under Sparse Data

Jul 2026 · 0 citations · 43 references
Mathematics

TL;DR

This paper provides a unified taxonomy and a large-scale, reproducible simulation benchmark; more than twenty tests are implemented in the open-source R package ebrahim.gof, and translate these findings into practical, evidence-based guidance for assessing logistic regression fit.

Abstract

Binary logistic regression is among the most widely used classification algorithms, yet a classifier is only trustworthy if its predicted probabilities are well calibrated. The classical checks -- the Pearson chi-square and deviance statistics -- break down precisely in the modern setting where predictors are continuous and the data are sparse (one covariate pattern per observation). Four decades of research have produced dozens of alternative goodness-of-fit and calibration algorithms, yet practitioners still default to the Hosmer-Lemeshow test because it ships with their software. This paper provides a unified taxonomy and a large-scale, reproducible simulation benchmark; more than twenty tests are implemented in the open-source R package ebrahim.gof. We evaluate them across five covariate distributions and four misspecification scenarios, with 10,000 replications each, measuring both Type I error and power. Several classical tests prove liberal, rejecting correct models far too often, while others have little power. A compact core -- McCullagh, Osius-Rojek, le Cessie-van Houwelingen, Stute-Zhu, and the GiViTI calibration test -- delivers the best balance of correct size and high power, and is consistently more powerful than the ubiquitous Hosmer-Lemeshow test. A low-birth-weight application reinforces the point: a model with omitted interactions slips past nearly every test, exposed only by pairing sensitive tests with a calibration (reliability) curve. We translate these findings into practical, evidence-based guidance for assessing logistic regression fit.

View source

Similar papers

Preprint Aug 2026

Goodness-of-Fit Tests and Calibration Machine-Learning Algorithms for Logistic Regression with Sparse Data

Assessing the goodness-of-fit of a logistic regression model is a critical prerequisite before the model is used for inference. However, goodness-of-fit (GOF) tests such as the chi-square and deviance tests often give invalid results when the data are"sparse"-- a common issue with continuous predictors like age or weig...

Ebrahim Khaled Ebrahim · 0 citations
Preprint Jul 2026

A directional Hosmer-Lemeshow goodness-of-fit test for sparse logistic regression

Goodness-of-fit assessment for the binary logistic regression model is difficult when covariates are continuous: the data are effectively sparse, the classical Pearson and deviance tests fail, and practitioners rely on partition-based tests, such as the Hosmer-Lemeshow test, that group observations before comparing obs...

Ebrahim Khaled Ebrahim, Ahmed El-Kotory · 1 citation
Open access Jul 2026

A Statistical Analysis of Classical and Advanced Regression Models in Data Science

Regression models remain foundational tools of both mathematical statistics and modern applied data science for prediction, explanation, and model comparison. In applied statistical modelling there is always a fundamental practical trade-off that must be considered. Classical regression models are attractive mainly bec...

Abdulwase Osmani, Abdul Qahar Majeedi · 0 citations
Preprint Aug 2026

Nonparametric Goodness-of-fit Testing under Covariate Shift

This paper develops procedures for nonparametric goodness-of-fit testing under covariate shift, where labelled data are drawn from a source population but goodness-of-fit is evaluated for a target population. The distribution mismatch is quantified by either a bounded moment condition or a sub-exponential tail conditio...

Zhengyou Hou, Dong Xia · 0 citations
Open access Aug 2026

Power Comparison of Selected Robust Non Parametric Tests for Assessing the Normality of Regression Residuals: A Monte Carlo Simulation Study

Assessing the normality of regression residuals is an important component of regression diagnostics because departures from normality can affect statistical inference, particularly in small and moderate samples. This study compared the empirical Type I error rates and empirical power of six selected normality tests for...

V. C. Ikwuka, C. H. Nwankwo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.