Evaluating Missing Data Imputation Strategies for Environmental Mixture Models: A Simulation Study and Applied Analysis
Missing data are common in environmental mixture studies and can bias inference if not properly addressed. This study evaluates how different imputation strategies influence variable selection performance in six mixture modeling frameworks: Weighted Quantile Sum (WQS), Bayesian WQS (BWQS), Quantile g-computation (Q-gcomp), Bayesian Kernel Machine Regression (BKMR), Elastic Net, and Least Absolute Shrinkage and Selection Operator (LASSO). Using Monte Carlo simulations (MCS), we generated multivariate normal (MVNORM) and multivariate t- (MVT) distributed exposures under linear and nonlinear outcome structures. Missingness was introduced at 25% for exposure and outcome variables, and at 5%–25% under both missing-at-random (MAR) and missing-not-at-random (MNAR) mechanisms. The study compares single imputation (SI) methods (mean, median, and k-nearest neighbors [KNN]) with multiple imputation approaches [MI] (MICE and Amelia) based on sensitivity (SE), specificity (SP), and false discovery rate (FDR). Across simulations, MI consistently improved variable selection performance under MAR, whereas SI and listwise deletion increased variability and reduced accuracy. Under MNAR, performance declined across all methods, with greater instability observed in flexible models such as BKMR and WQS. Q-gcomp demonstrated the most consistent balance across SE, SP, and FDR and remained relatively robust to violations of the MAR assumption. Additional analyses highlight the importance of imputation model specification: excluding relevant covariates (e.g. income) from the imputation process increased variability and reduced stability, particularly for flexible models and binary outcomes. Application to NHANES 2007–2014 data (n = 8,233) showed that MI improved the stability of variable importance estimates, with BWQS and Q-gcomp yielding more reproducible exposure rankings. Overall, results demonstrate that both the missingness mechanism and the choice and specification of imputation strategy substantially influence variable selection in mixture models, underscoring the importance of carefully designed imputation procedures in environmental epidemiology.