Skip to content

Data Augmentation for Machine Learning on Missing Data

Aug 2026 · Programming and computer software · Vol 52, pp. 344 - 357 · 0 citations · 40 references
Computer Science

TL;DR

It is proved that drop-and-replace augmentation allows one to maximize balanced accuracy when learning on small linearly separable datasets containing missing values, and is more efficient when learning on missing data.

View source

Similar papers

Open access 2026

Collection: Benchmark Datasets for the Evaluation of Data Imputation Algorithms

Complete and labeled datasets were gathered from the web and processed to simulate missing completely at random (MCAR) and missing at random (MAR) missingness mechanisms, which are inspired by common missing-data patterns. From the original 12 datasets, we generated 2400 benchmark datasets with missing data using the M...

Alice B. Nogueira, I. Carvalho · 0 citations
Open access Aug 2026

Machine Learning-Based Imputation for Breast Cancer Prediction: Evaluating Performance Under Complex Missing Data Mechanisms

This study systematically compared statistical and machine learning-based imputation methods using two publicly available breast cancer datasets representing complementary clinical settings to highlight the importance of considering dataset characteristics, missing-data mechanisms, and the intended analytical objective...

Nyatuga Gideon Nyakundi, John Ndiritu, Ivivi J. Mwaniki et al. · 0 citations
Open access Aug 2026

COMPARATIVE ANALYSIS OF RANDOM FOREST AND SUPPORT VECTOR MACHINE ALGORITHMS FOR DIABETES MELLITUS PREDICTION

Comparisons of the performance of the Random Forest and Support Vector Machine algorithms in predicting diabetes and the effect of applying the Synthetic Minority Over-sampling Technique to imbalanced data show that Random Forest outperforms SVM.

Baharudin Yusuf · 0 citations
Open access 2026

Optimizing diabetes prediction in machine learning models: Evaluating the effectiveness of a novel class imbalance technique—adaptive synthetic class balancing with class proportion filtering

The skewedness of results when predicting diabetes is mostly due to uneven distribution of data, especially in reducing detection rates of real patients. These are errors which cause delay in treatment or incorrect diagnosis. This work suggests a counter plan to this assumption, which is Adaptive Synthetic Class Balanc...

Pankaj Beldar, Snehal M. Kamalapur, Priti Vaidya et al. · 0 citations
Open access Sep 2026

Comparing the use of supervised machine learning variable selection methods in the context of two-group classification in the psychological and health sciences

Introduction Variable selection (VS) is crucial for building accurate and generalizable classification models. Reducing the necessary number of variables improves model efficiency, interpretability, and generalizability while reducing data collection burden. Despite the availability of various VS methods, their compara...

Catherine M. Bain, Ding-Jing Shi, Y. Banad et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.