Aug 2026· Programming and computer software· Vol 52, pp. 344 - 357· 0 citations· 40 references
Computer Science
TL;DR
It is proved that drop-and-replace augmentation allows one to maximize balanced accuracy when learning on small linearly separable datasets containing missing values, and is more efficient when learning on missing data.
Complete and labeled datasets were gathered from the web and processed to simulate missing completely at random (MCAR) and missing at random (MAR) missingness mechanisms, which are inspired by common missing-data patterns. From the original 12 datasets, we generated 2400 benchmark datasets with missing data using the M...
Alice B. Nogueira, I. Carvalho· IEEE Data Descriptions· 0 citations
This study systematically compared statistical and machine learning-based imputation methods using two publicly available breast cancer datasets representing complementary clinical settings to highlight the importance of considering dataset characteristics, missing-data mechanisms, and the intended analytical objective...
Nyatuga Gideon Nyakundi, John Ndiritu, Ivivi J. Mwaniki et al.· AppliedMath· 0 citations
Comparisons of the performance of the Random Forest and Support Vector Machine algorithms in predicting diabetes and the effect of applying the Synthetic Minority Over-sampling Technique to imbalanced data show that Random Forest outperforms SVM.
Baharudin Yusuf· Jurnal Informatika dan Tekni...· 0 citations
The skewedness of results when predicting diabetes is mostly due to uneven distribution of data, especially in reducing detection rates of real patients. These are errors which cause delay in treatment or incorrect diagnosis. This work suggests a counter plan to this assumption, which is Adaptive Synthetic Class Balanc...
Pankaj Beldar, Snehal M. Kamalapur, Priti Vaidya et al.· Sigma Journal of Engineering...· 0 citations
Introduction Variable selection (VS) is crucial for building accurate and generalizable classification models. Reducing the necessary number of variables improves model efficiency, interpretability, and generalizability while reducing data collection burden. Despite the availability of various VS methods, their compara...
Catherine M. Bain, Ding-Jing Shi, Y. Banad et al.· Frontiers in Psychology· 0 citations
It is shown that preserving data quality should take priority over quantity when the missing rates are exceptionally high, providing the practical guidance for robust preprocessing design in federated learning systems.
Xiao Liu· Mathematical Modeling and Al...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.