Aug 2026· Applied Sciences· 0 citations· 50 references
TL;DR
A systematic comparison of backpropagation and ten variants of a genetic algorithm for training multi-layer perceptrons (MLPs), with particular focus on the role of crossover operators, helps clarify when gradient-free training is competitive and which evolutionary operators drive its effectiveness.
Abstract
Training neural networks with gradient-based methods such as backpropagation is the dominant paradigm, but it depends on differentiable loss functions and is sensitive to initialization and local minima. Evolutionary algorithms offer a gradient-free alternative, yet the influence of their internal operators on training quality remains insufficiently characterized. This study presents a systematic comparison of backpropagation and ten variants of a genetic algorithm (GA) for training multi-layer perceptrons (MLPs), with particular focus on the role of crossover operators. The evaluation covers four MLP architectures and ten classification datasets from the UCI Machine Learning Repository, differing in sample size, dimensionality, and number of classes. Each configuration was assessed using stratified 4-fold cross-validation with 30 independent repetitions, and accuracy served as the primary performance metric, with macro-F1 reported to assess classifier behavior on class-imbalanced datasets. Backpropagation achieved higher mean accuracy than every GA variant on nine of the ten datasets, with the largest margins on high-dimensional problems. The genetic algorithm proved competitive on simpler, class-balanced datasets, where its better-performing variants matched the gradient-based baseline within one to two percentage points, and, on the Heart disease dataset, every GA variant reached a higher mean accuracy than backpropagation across all four architectures, though absolute performance remained modest on this five-class problem. Among crossover operators, BLX-α and BLX-α-β combined with tournament selection and a high crossover probability yielded the strongest configurations, while averaging crossover performed worst, as it restricts offspring to the midpoint of the parents and cannot explore beyond the range already present in the population. Tournament selection consistently led to higher mean accuracy than roulette-wheel selection, and shallow but moderately wide architectures, which encode fewer trainable parameters and thus a shorter chromosome, proved more amenable to evolutionary training than the two-layer alternative. These findings clarify when gradient-free training is competitive and which evolutionary operators drive its effectiveness.
An improved multi-operator differential evolution algorithm, called NN-IMODE, tailored for comprehensive Feedforward Neural Networks optimization, encompassing the number of hidden layers, neurons per layer, weights, biases, and activation functions, is introduced.
Aridj Ferhat, Farouq Zitouni, Karam M. Sallam et al.· Cluster Computing· 0 citations
Results indicate that coupling high-order Runge-Kutta numerical integration with population-based stochastic search offers an effective, gradient-free strategy for MLP training and other nonlinear optimization problems.
Hisham M. Khudhur, Basma T. Fathy, R. Ramo· Journal of Intelligent &...· 0 citations
This paper presents a multi-stage evolutionary technique based on genetic algorithms for the effective training of RBF networks, applied to a large set of classification and data-fitting problems, yielding excellent results.
Ioannis G. Tsoulos, Vasileios Charilogis, Dimitrios G. Tsalikakis· Mathematics· 0 citations
This paper deeply integrates convex optimization theory with the backpropagation algorithm and constructs a novel stable and efficient training mechanism for neural networks that achieves favorable adaptability to both shallow fully connected networks and deep convolutional networks.
Weiwei Guo· Applied and Computational En...· 0 citations
The study shows how optimization techniques influence neural network training behavior and it combines mathematical optimization theory with real model performance.
Pu Gong· Applied and Computational En...· 0 citations
SVM was the most consistent and reliable method across the tested scenarios, achieving high classification accuracy over a wide range of training sample sizes and maintaining strong robustness under noisy labels, which makes it a strong baseline choice when reference data are limited or imperfect.
P. Kupidura, Michał Szkibiel· Reports on Geodesy and Geoin...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.