Skip to content
Open access

Eleven quick tips to reduce overfitting in machine learning

Aug 2026 · International Journal of Data Science and Analysis · Vol 22 · 0 citations · 139 references

TL;DR

Eleven practical tips for reducing unintentional overfitting in supervised biomedical machine learning studies are presented and stress principled data splitting, domain-informed preprocessing, controlled model complexity, systematic tuning, comprehensive performance evaluation, and robustness analysis.

Abstract

Overfitting is the excessive adaptation of a machine learning model to its training data and remains a persistent challenge in biomedical informatics. In supervised learning, models may capture patterns “too well,” failing to generalize to unseen data and yielding overly optimistic performance estimates. The problem is especially acute in biomedical settings, where datasets are high-dimensional, heterogeneous, and often of few samples. Numerous strategies have been proposed to mitigate overfitting in bioinformatics and health informatics. However, even established techniques can produce misleading results if misapplied, for example through data leakage, excessive hyperparameter tuning, inappropriate preprocessing, or inadequate validation. To address these pitfalls, we present eleven practical tips for reducing unintentional overfitting in supervised biomedical machine learning studies. The recommendations stress principled data splitting, domain-informed preprocessing, controlled model complexity, systematic tuning, comprehensive performance evaluation, and robustness analysis. Rather than offering an exhaustive treatment, we provide an accessible, practice-oriented guide to support more reliable and reproducible machine learning research. Although developed for biomedical informatics, these quick tips are broadly applicable across disciplines using supervised machine learning.

Read PDF

Similar papers

Open access Aug 2026

Eleven quick tips for Biomedical Federated Learning

Ten tips for successfully and sustainably implementing Federated learning for Biomedical applications, ensuring both ethical data governance and improved model performance in sensitive domains are outlined.

Kyle Ellrott, V. Malladi, J. Bélisle-Pipon et al. · 0 citations
Aug 2026

Post-pretrained lasso statistical inference

This work proposes post-pretrained lasso selective inference (PPL-SI), a novel selective inference method designed to provide statistically valid p values for the pretrained lasso that reliably controls false discoveries and significantly improves the detection of biologically relevant features compared to traditional...

Cao Huyen My, Nguyen Vu Khai Tam, Vo Nguyen Le Duy · 0 citations
Open access 2026

Recommendation: How to Use Synthetic Data in Machine Learning or Decision Support

The impact of balancing real and SD is examined and some recommendations that researchers can further utilise to improve the ML model’s training process are provided and which approaches to adopt are considered.

Majid Liaquat, Chris D. Nugent, I. Cleland et al. · 0 citations
Open access 2026

An Interpretable XGBoost Model for Diabetes Prediction: Nested Cross-Validation, Calibration, and SHAP Analysis

Under a leakage-controlled, unbiased evaluation, XGBoost provided moderate but trustworthy discrimination together with well-calibrated probabilities for diabetes prediction, while SHAP confirmed clinically plausible predictors.

Z. Kucukakcali, I. Cicek · 0 citations
Open access Aug 2026

Can GPT Be Used as an Alternative Prediction Model to Traditional Machine Learning and Neural Networks on Low-Volume Clinical Data?

The proposed GPT2-based table-to-text framework provides a practical and clinically interpretable approach for disease prediction from limited structured healthcare data and demonstrates strong potential for early risk detection, transparent clinical decision support, and reliable deployment in real-world low-resource...

S. Bin Akter, S. Akter, D. Eisenberg et al. · 0 citations
Open access Aug 2026

Artificial Intelligence-Based Clinical Risk Prediction Using Ensemble Machine Learning Algorithms

The outcomes demonstrate that the suggested ANN architecture can accurately and consistently assess clinical risk, which in turn allows for the early identification of high-risk patients and aids healthcare providers in making data-informed therapeutic decisions.

Shivani Jain · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.