Jul 2026· Reports on Geodesy and Geoinformatics· Vol 122, pp. 1 - 13· 0 citations· 37 references
TL;DR
SVM was the most consistent and reliable method across the tested scenarios, achieving high classification accuracy over a wide range of training sample sizes and maintaining strong robustness under noisy labels, which makes it a strong baseline choice when reference data are limited or imperfect.
Abstract
Abstract This study evaluates the performance of selected machine learning methods, Maximum Likelihood (MLC), Random Forest (RF), Extreme Gradient Boosting (XGB), Support Vector Machine (SVM), and Artificial Neural Networks (ANN), for land use/land cover (LULC) classification using Sentinel-2 satellite imagery. Each algorithm was tested across multiple classification scenarios that systematically varied both training sample size and training sample quality. To emulate realistic imperfections in reference data, controlled levels of label noise were introduced into the training set, and the resulting changes in classification performance were analysed. In addition, each model’s susceptibility to overfitting was assessed by comparing performance on training data and independent test data. To reduce the influence of a single random draw of training and testing pixels, all sampling-based experiments were repeated 10 times using different random seed values, and the reported results were summarized using repeated-run statistics. This design enabled a comparable assessment of how classification accuracy, stability, and generalization depend on dataset size and quality. The results indicate that SVM was the most consistent and reliable method across the tested scenarios, achieving high classification accuracy over a wide range of training sample sizes and maintaining strong robustness under noisy labels. Compared with the other methods, SVM showed lower sensitivity to degraded training data and a smaller tendency to overfit, which makes it a strong baseline choice when reference data are limited or imperfect.
Introduction Variable selection (VS) is crucial for building accurate and generalizable classification models. Reducing the necessary number of variables improves model efficiency, interpretability, and generalizability while reducing data collection burden. Despite the availability of various VS methods, their compara...
Catherine M. Bain, Ding-Jing Shi, Y. Banad et al.· Frontiers in Psychology· 0 citations
Overall, logistic regression demonstrates strong and good generalization ability, while ensemble methods constitute robust alternatives for binary classification on structured data.
Zahra Benider, H. Bouzahir, Jaafar Idrais· EPJ Web of Conferences· 0 citations
Class imbalance remains a critical challenge in supervised learning, often biasing classifiers toward majority classes. While resampling techniques like Synthetic Minority Oversampling Technique (SMOTE) are widely used, the combined effect of data balancing and hyperparameter optimization across diverse datasets is rar...
Necati Vardar, Mehmet Fatih Ören· Sakarya University Journal o...· 0 citations
This study examined advanced machine learning approaches for predictive analytics by comparatively evaluating the classification performance of Support Vector Machine (SVM), Random Forest, and Gradient Boosting in high-dimensional data. A high-dimensional dataset containing multiple observations and a large number of f...
Imtiaz Ali, M. Ali, Shahrukh Nawaz· Journal of Global Social Tra...· 0 citations
Tree-based models exhibited the largest reductions in accuracy and the steepest increases in calibration error under scarcity, whereas logistic regression preserved threshold balance and probability calibration at negligible computational cost.
Jairo Devon A. Daquioag, E. Ayo· Engineering and Technology J...· 0 citations
Class imbalance is a common problem in machine learning classification tasks and may negatively affect the reliability and predictive performance of classification models. The research object is the comprehensive evaluation of different resampling techniques applied to a publicly available breast cancer dataset consist...
Ogtay Safaraliyev, K. Mammadova· EUREKA Physics and Engineeri...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.