Jul 2026· Applied and Computational Engineering· Vol 254, pp. 59-66· 0 citations
TL;DR
The study shows how optimization techniques influence neural network training behavior and it combines mathematical optimization theory with real model performance.
Abstract
Neural networks are often employed in artificial intelligence because they are capable of learning patterns and approximating nonlinear functions from input data. But the performance of a neural network is not only decided by the model structure, but also by the optimization technique utilized in the training process. In this research, we compare full-batch Gradient Descent versus mini-batch Stochastic Gradient Descent in a neural network function approximation problem. A tiny feedforward neural network is trained to approximate the nonlinear function. The experiment is implemented in Python and the comparison is made based on training loss curves, final training means squared error, final test mean squared error and prediction visualization. The results reveal that the full-batch Gradient Descent gives a smoother loss curve since in each update the entire training dataset is used. Mini-batch Stochastic Gradient Descent on the other hand shows more noticeable oscillations but reaches a lower ultimate training MSE and test MSE in this experimental context. This implies that stochastic updates can be less stable at each step but can still assist the model in achieving greater approximation performance in the same number of epochs. The study shows how optimization techniques influence neural network training behavior and it combines mathematical optimization theory with real model performance.
A broad class of neural networks trained under the $\ell^2$ loss by gradient descent with weight decay with weight decay is studied, and the convergence of GD to a neighbourhood of the global minimizers of the empirical loss is proved.
Yu-Qing Wang, I. Kevrekidis, Mikhail Belkin· 1 citation
Estimation of a regression function from independent and identically distributed data is considered. The $L_2$ error with integration with respect to the design variable is used as the error criterion. An initially randomly pruned fully connected deep neural network with logistic squasher as activation function is fitt...
M. Kohler, Vincent Molinero Römer, Adam Krzyżak· 0 citations
We establish the first convergence guarantees of gradient descent for general feedforward neural networks of any width or depth, with any initialization or dataset. We only assume that the activation functions are linearly bounded, Lipschitz continuous, and Lipschitz smooth---properties that hold for linear, tanh, soft...
An improved multi-operator differential evolution algorithm, called NN-IMODE, tailored for comprehensive Feedforward Neural Networks optimization, encompassing the number of hidden layers, neurons per layer, weights, biases, and activation functions, is introduced.
Aridj Ferhat, Farouq Zitouni, Karam M. Sallam et al.· Cluster Computing· 0 citations
A systematic comparison of backpropagation and ten variants of a genetic algorithm for training multi-layer perceptrons (MLPs), with particular focus on the role of crossover operators, helps clarify when gradient-free training is competitive and which evolutionary operators drive its effectiveness.
Mikołaj Petecki, W. Książek, Artur Niewiarowski· Applied Sciences· 0 citations
Accurate nonlinear modeling underpins every layer of modern engineering. Sampling neural network (SNN) is a new fitting network, which has a simple structure, clear physical concept, concise algorithm, stable performance, and adopts a new error diffusion training method similar to the diffusion of neural stimuli in liv...
Ling-Yan Wu, Gang Cai· Electrica· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.