Skip to content
Open access

A comparison of missing data approaches for linear regression with missing not at random outcome and predictors

Jul 2026 · Statistical Methods in Medical Research · Vol 35, pp. 1658 - 1677 · 0 citations · 46 references
Medicine

Abstract

While most methods for missing not at random (MNAR) data in regression models address MNAR outcomes assuming fully observed predictors, real-world observational health and longitudinal studies often violate this assumption. This paper compares several approaches for handling MNAR data in linear regression when missingness depends on both partially observed outcomes and predictors. Through extensive simulations, we evaluate complete-case analysis, multiple imputation assuming missing at random, maximum likelihood estimation via the Heckman selection model, uncertainty intervals, multiple imputation under the Heckman selection model, not-at-random fully conditional specification, imputation stacking, and random indicator imputation. None of the methods consistently produced unbiased estimates or nominal coverage across all scenarios. However, not-at-random fully conditional specification was straightforward to implement and yielded coverage close to the nominal level in most scenarios, provided that the sensitivity parameters were specified near their true values. Our results highlight the importance of sensitivity analyses exploring various full-data models and careful parameter specification when addressing MNAR in both outcomes and predictors. We illustrate such sensitivity analyses using Betula study data on the relationship between longitudinal memory change and grey matter volume in aging. The association remained significant across most considered MNAR scenarios, reinforcing existing evidence for this relationship.

Read PDF