Skip to content
Review Open access

A Comprehensive Review on Missing Data Imputation Techniques

2026 · ITEGAM- Journal of Engineering and Technology for Industrial Applications (ITEGAM-JETIA) · 0 citations

Abstract

Missing data is a fundamental challenge in scientific research and often leads to biased results and weakened statistical inference. To address this, imputation methods are essential to preserve the original information. This study compares a wide range of imputation techniques categorized into two main types: traditional statistical methods (mean, median, mode, last observation carried forward, regression, multiple imputation), and machine learning and deep learning methods including algorithms such as (k-nearest neighbors, random forests, support vector machines, neural networks, generative adversarial networks). The results indicate that machine learning algorithms, particularly missForest and KNN, consistently provide high accuracy by modeling complex and nonlinear relationships. Furthermore, GAN-based methods such as GAIN, GIMIN, and MisGIMIN represent significant advancements, especially for high-dimensional data with missing rates exceeding 80%. Specifically, the GIMIN algorithm demonstrated superior performance in root mean square error (RMSE) accuracy at high levels of missingness, while MisGAN achieved the lowest FID error. The study emphasizes selecting methods based on the characteristics of the dataset and recommends advanced machine learning algorithms to ensure unbiased inference, while cautioning against simple imputation techniques that may introduce bias.

Read PDF