Skip to content
Open access

Evaluating machine learning algorithms: The role of sample size and quality

Jul 2026 · Reports on Geodesy and Geoinformatics · Vol 122, pp. 1 - 13 · 0 citations · 37 references

TL;DR

SVM was the most consistent and reliable method across the tested scenarios, achieving high classification accuracy over a wide range of training sample sizes and maintaining strong robustness under noisy labels, which makes it a strong baseline choice when reference data are limited or imperfect.

Abstract

Abstract This study evaluates the performance of selected machine learning methods, Maximum Likelihood (MLC), Random Forest (RF), Extreme Gradient Boosting (XGB), Support Vector Machine (SVM), and Artificial Neural Networks (ANN), for land use/land cover (LULC) classification using Sentinel-2 satellite imagery. Each algorithm was tested across multiple classification scenarios that systematically varied both training sample size and training sample quality. To emulate realistic imperfections in reference data, controlled levels of label noise were introduced into the training set, and the resulting changes in classification performance were analysed. In addition, each model’s susceptibility to overfitting was assessed by comparing performance on training data and independent test data. To reduce the influence of a single random draw of training and testing pixels, all sampling-based experiments were repeated 10 times using different random seed values, and the reported results were summarized using repeated-run statistics. This design enabled a comparable assessment of how classification accuracy, stability, and generalization depend on dataset size and quality. The results indicate that SVM was the most consistent and reliable method across the tested scenarios, achieving high classification accuracy over a wide range of training sample sizes and maintaining strong robustness under noisy labels. Compared with the other methods, SVM showed lower sensitivity to degraded training data and a smaller tendency to overfit, which makes it a strong baseline choice when reference data are limited or imperfect.

Read PDF

Similar papers

Open access Sep 2026

Comparing the use of supervised machine learning variable selection methods in the context of two-group classification in the psychological and health sciences

Introduction Variable selection (VS) is crucial for building accurate and generalizable classification models. Reducing the necessary number of variables improves model efficiency, interpretability, and generalizability while reducing data collection burden. Despite the availability of various VS methods, their compara...

Catherine M. Bain, Ding-Jing Shi, Y. Banad et al. · 0 citations
Conference Open access 2026

Supervised Learning for Classification in Data Science: A Comparative Perspective

Overall, logistic regression demonstrates strong and good generalization ability, while ensemble methods constitute robust alternatives for binary classification on structured data.

Zahra Benider, H. Bouzahir, Jaafar Idrais · 0 citations
Open access Sep 2026

A Comparative Performance Evaluation of Classification Algorithms on Imbalanced Datasets

Class imbalance remains a critical challenge in supervised learning, often biasing classifiers toward majority classes. While resampling techniques like Synthetic Minority Oversampling Technique (SMOTE) are widely used, the combined effect of data balancing and hyperparameter optimization across diverse datasets is rar...

Necati Vardar, Mehmet Fatih Ören · 0 citations
Open access Sep 2026

Advanced Machine Learning Approaches for Predictive Analytics: Comparative Evaluation of Classification Performance in High-Dimensional Data

This study examined advanced machine learning approaches for predictive analytics by comparatively evaluating the classification performance of Support Vector Machine (SVM), Random Forest, and Gradient Boosting in high-dimensional data. A high-dimensional dataset containing multiple observations and a large number of f...

Imtiaz Ali, M. Ali, Shahrukh Nawaz · 0 citations
#software testing Open access Nov 2026

How Classification Models Degrade as Data Diminish: A Multi-Criteria Analysis of Tabular Classifiers Under Controlled Data Scarcity

Tree-based models exhibited the largest reductions in accuracy and the steepest increases in calibration error under scarcity, whereas logistic regression preserved threshold balance and probability calibration at negligible computational cost.

Jairo Devon A. Daquioag, E. Ayo · 0 citations
Open access Sep 2026

Comparative evaluation of resampling techniques for improving the performance of classification algorithms on a breast cancer dataset

Class imbalance is a common problem in machine learning classification tasks and may negatively affect the reliability and predictive performance of classification models. The research object is the comprehensive evaluation of different resampling techniques applied to a publicly available breast cancer dataset consist...

Ogtay Safaraliyev, K. Mammadova · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.