Skip to content
Open access

Not All Missing Data are Equal: Choosing the Right Imputation Method for Binary Datasets

Jul 2026 · Quality and Reliability Engineering International · 0 citations · 23 references

Abstract

Missing binary predictors are common in reliability, quality control, and industrial decision systems, yet imputation methods are often chosen by convenience rather than evidence. We conduct a Monte Carlo study comparing mode substitution, sequential hot‐deck, missForest, MICE, and KNN with three neighbourhood sizes under MCAR, MAR, and MNAR missingness, across missingness rates from 5% to 50% and two predictor‐dependence structures. Performance is evaluated on three targets: exact recovery of missing binary cells, recovery of logistic‐regression coefficients, and downstream classification using logistic regression, naive Bayes, support vector machines, and random forests. The results reveal a clear trade‐off. KNN is strongest for exact cell recovery under MCAR and MAR, whereas missForest performs best under MNAR. MICE is the most reliable choice for downstream predictive performance across learners and missingness mechanisms. By contrast, mode imputation and sequential hot‐deck achieve the best coefficient recovery. The main implication is operational: in binary‐data environments, imputation should be chosen to match the analytical objective–reconstruction, inference, or prediction–because no single method dominates all targets simultaneously.

Read PDF

Similar papers

Open access Jul 2026

A comparison of missing data approaches for linear regression with missing not at random outcome and predictors

While most methods for missing not at random (MNAR) data in regression models address MNAR outcomes assuming fully observed predictors, real-world observational health and longitudinal studies often violate this assumption. This paper compares several approaches for handling MNAR data in linear regression when missingness depends on both partially observed outcomes and predictors. Through extensive simulations, we evaluate complete-case analysis, multiple imputation assuming missing at random, maximum likelihood estimation via the Heckman selection model, uncertainty intervals, multiple imputation under the Heckman selection model, not-at-random fully conditional specification, imputation stacking, and random indicator imputation. None of the methods consistently produced unbiased estimates or nominal coverage across all scenarios. However, not-at-random fully conditional specification was straightforward to implement and yielded coverage close to the nominal level in most scenarios, provided that the sensitivity parameters were specified near their true values. Our results highlight the importance of sensitivity analyses exploring various full-data models and careful parameter specification when addressing MNAR in both outcomes and predictors. We illustrate such sensitivity analyses using Betula study data on the relationship between longitudinal memory change and grey matter volume in aging. The association remained significant across most considered MNAR scenarios, reinforcing existing evidence for this relationship.

T. Gorbach, Tim P. Morris, James E. Carpenter · 0 citations
Review Open access 2026

A Comprehensive Review on Missing Data Imputation Techniques

Missing data is a fundamental challenge in scientific research and often leads to biased results and weakened statistical inference. To address this, imputation methods are essential to preserve the original information. This study compares a wide range of imputation techniques categorized into two main types: traditional statistical methods (mean, median, mode, last observation carried forward, regression, multiple imputation), and machine learning and deep learning methods including algorithms such as (k-nearest neighbors, random forests, support vector machines, neural networks, generative adversarial networks). The results indicate that machine learning algorithms, particularly missForest and KNN, consistently provide high accuracy by modeling complex and nonlinear relationships. Furthermore, GAN-based methods such as GAIN, GIMIN, and MisGIMIN represent significant advancements, especially for high-dimensional data with missing rates exceeding 80%. Specifically, the GIMIN algorithm demonstrated superior performance in root mean square error (RMSE) accuracy at high levels of missingness, while MisGAN achieved the lowest FID error. The study emphasizes selecting methods based on the characteristics of the dataset and recommends advanced machine learning algorithms to ensure unbiased inference, while cautioning against simple imputation techniques that may introduce bias.

NoorAldeen Salam Taher, N. J. Mohammed · 0 citations
Open access Aug 2026

Machine Learning-Based Imputation for Breast Cancer Prediction: Evaluating Performance Under Complex Missing Data Mechanisms

Missing data remain a major challenge in breast cancer research because they can introduce bias, reduce statistical efficiency, and compromise the performance of predictive models. Although numerous imputation techniques have been proposed, their comparative performance under different missing-data mechanisms and their impact on downstream classification remain inadequately understood. This study systematically compared statistical and machine learning-based imputation methods using two publicly available breast cancer datasets representing complementary clinical settings. The methods were evaluated under simulated Missing Completely at Random (MCAR), Missing at Random (MAR), and Missing Not at Random (MNAR) mechanisms using both reconstruction accuracy and downstream classification performance. The results showed that no single imputation method consistently achieved the best performance across both datasets. Regularized regression and machine learning-based methods generally outperformed conventional statistical approaches, although the optimal method depended on the characteristics of the dataset. Furthermore, the best-performing imputation methods preserved downstream classification performance despite the introduction of missing data, demonstrating that reconstruction accuracy alone is insufficient for selecting imputation strategies intended for predictive modelling. Overall, the findings highlight the importance of considering dataset characteristics, missing-data mechanisms, and the intended analytical objective when selecting imputation methods. The proposed evaluation framework provides a robust approach for assessing missing-data handling strategies in breast cancer prediction studies and other biomedical machine learning applications.

Nyatuga Gideon Nyakundi, John Ndiritu, Ivivi J. Mwaniki et al. · 0 citations
Open access Jul 2026

Evaluating Missing Data Imputation Strategies for Environmental Mixture Models: A Simulation Study and Applied Analysis

Missing data are common in environmental mixture studies and can bias inference if not properly addressed. This study evaluates how different imputation strategies influence variable selection performance in six mixture modeling frameworks: Weighted Quantile Sum (WQS), Bayesian WQS (BWQS), Quantile g-computation (Q-gcomp), Bayesian Kernel Machine Regression (BKMR), Elastic Net, and Least Absolute Shrinkage and Selection Operator (LASSO). Using Monte Carlo simulations (MCS), we generated multivariate normal (MVNORM) and multivariate t- (MVT) distributed exposures under linear and nonlinear outcome structures. Missingness was introduced at 25% for exposure and outcome variables, and at 5%–25% under both missing-at-random (MAR) and missing-not-at-random (MNAR) mechanisms. The study compares single imputation (SI) methods (mean, median, and k-nearest neighbors [KNN]) with multiple imputation approaches [MI] (MICE and Amelia) based on sensitivity (SE), specificity (SP), and false discovery rate (FDR). Across simulations, MI consistently improved variable selection performance under MAR, whereas SI and listwise deletion increased variability and reduced accuracy. Under MNAR, performance declined across all methods, with greater instability observed in flexible models such as BKMR and WQS. Q-gcomp demonstrated the most consistent balance across SE, SP, and FDR and remained relatively robust to violations of the MAR assumption. Additional analyses highlight the importance of imputation model specification: excluding relevant covariates (e.g. income) from the imputation process increased variability and reduced stability, particularly for flexible models and binary outcomes. Application to NHANES 2007–2014 data (n = 8,233) showed that MI improved the stability of variable importance estimates, with BWQS and Q-gcomp yielding more reproducible exposure rankings. Overall, results demonstrate that both the missingness mechanism and the choice and specification of imputation strategy substantially influence variable selection in mixture models, underscoring the importance of carefully designed imputation procedures in environmental epidemiology.

Y. S. Boafo, S. Mostafa, E. Obeng-Gyasi · 0 citations
Open access Aug 2026

Columnwise neural imputation for incomplete ordinal psychometric data.

Missing data are pervasive in psychological and educational assessments. Naive remedies, including listwise deletion and item-mean imputation, often degrade research validity and misinform subsequent decisions. Recent advances in artificial neural networks have demonstrated their efficacy in prediction-related tasks by using observed features to infer unknown values. Building on this potential, we propose the columnwise neural imputation (COLNI) algorithm to impute missing ordinal responses in psychometric data. Simulation studies demonstrated that, when benchmarked against conventional methods, COLNI more accurately recovered item means, inter-item correlations, and person and item parameters under the multidimensional graded response model. We further evaluated COLNI using data from the Short Dark Triad test, confirming its effectiveness in a multidimensional empirical setting. We conclude with implementation guidelines and avenues for refining and extending this artificial neural network-based imputation approach in future research. (PsycInfo Database Record (c) 2026 APA, all rights reserved).

Longfei Zhang, Minjeong Jeon, Ping Chen · 0 citations
Review Open access Aug 2026

Bayesian model comparison for random effects probit models with missing covariates

For model comparison in random effects probit models with incompletely observed covariates, this paper develops a Bayesian data-augmentation workflow in which latent Gaussian responses, random effects, and missing covariate values are updated within a common augmented sampling scheme. Because specifying a fully parametric joint model for mixed continuous and categorical covariates is often unattractive in survey applications, missing covariates are updated by a decision-tree-assisted Bayesian-bootstrap step. Competing models are evaluated conditionally on one common medoid completion using Chib’s method with reduced Gibbs sampling; sensitivity is assessed with respect to the Chib evaluation point, the regression-coefficient prior, and the medoid reference model. The simulation study compares the proposed approach with complete case analysis, multiple imputation by chained equations, missForest single imputation, information criteria, and predictive criteria under MCAR, cross-dependent MAR-type, and self-masked MNAR scenarios. An empirical illustration based on the National Educational Panel Study demonstrates how the method can be used for comparing labor-market models of current employment when competence measures and employment-history covariates are incompletely observed. The results show that missing covariates can materially affect model rankings, and that the proposed workflow provides a transparent evidence-based comparison of nested and non-nested random effects probit specifications under incomplete covariate information.

Michael Bergrab · 0 citations