Skip to content
Open access

A Comparative Study of Gradient Descent and Stochastic Gradient Descent in Neural Network Function Approximation

Jul 2026 · Applied and Computational Engineering · Vol 254, pp. 59-66 · 0 citations

TL;DR

The study shows how optimization techniques influence neural network training behavior and it combines mathematical optimization theory with real model performance.

Abstract

Neural networks are often employed in artificial intelligence because they are capable of learning patterns and approximating nonlinear functions from input data. But the performance of a neural network is not only decided by the model structure, but also by the optimization technique utilized in the training process. In this research, we compare full-batch Gradient Descent versus mini-batch Stochastic Gradient Descent in a neural network function approximation problem. A tiny feedforward neural network is trained to approximate the nonlinear function. The experiment is implemented in Python and the comparison is made based on training loss curves, final training means squared error, final test mean squared error and prediction visualization. The results reveal that the full-batch Gradient Descent gives a smoother loss curve since in each update the entire training dataset is used. Mini-batch Stochastic Gradient Descent on the other hand shows more noticeable oscillations but reaches a lower ultimate training MSE and test MSE in this experimental context. This implies that stochastic updates can be less stable at each step but can still assist the model in achieving greater approximation performance in the same number of epochs. The study shows how optimization techniques influence neural network training behavior and it combines mathematical optimization theory with real model performance.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

A Theoretical Analysis of Generalization Dynamics in Neural Networks under Gradient Descent with Weight Decay

A broad class of neural networks trained under the $\ell^2$ loss by gradient descent with weight decay with weight decay is studied, and the convergence of GD to a neighbourhood of the global minimizers of the empirical loss is proved.

Yu-Qing Wang, I. Kevrekidis, Mikhail Belkin · 1 citation
Preprint Aug 2026

Learning of deep neural network regression estimates using gradient descent with pruning

Estimation of a regression function from independent and identically distributed data is considered. The $L_2$ error with integration with respect to the design variable is used as the error criterion. An initially randomly pruned fully connected deep neural network with logistic squasher as activation function is fitt...

M. Kohler, Vincent Molinero Römer, Adam Krzyżak · 0 citations
#machine learning Preprint Aug 2026

Convergence Guarantees of Gradient Descent for Neural Networks via Generalized Lipschitz Smoothness

We establish the first convergence guarantees of gradient descent for general feedforward neural networks of any width or depth, with any initialization or dataset. We only assume that the activation functions are linearly bounded, Lipschitz continuous, and Lipschitz smooth---properties that hold for linear, tanh, soft...

Siqiao Mu, Diego Klabjan · 0 citations
Sep 2026

NN-IMODE: a novel approach for simultaneous optimization of architecture and parameters in feedforward neural networks

An improved multi-operator differential evolution algorithm, called NN-IMODE, tailored for comprehensive Feedforward Neural Networks optimization, encompassing the number of hidden layers, neurons per layer, weights, biases, and activation functions, is introduced.

Aridj Ferhat, Farouq Zitouni, Karam M. Sallam et al. · 0 citations
Open access Aug 2026

Evolutionary Training of Neural Networks: The Role of Crossover Operators in Genetic Algorithms Compared with Backpropagation

A systematic comparison of backpropagation and ten variants of a genetic algorithm for training multi-layer perceptrons (MLPs), with particular focus on the role of crossover operators, helps clarify when gradient-free training is competitive and which evolutionary operators drive its effectiveness.

Mikołaj Petecki, W. Książek, Artur Niewiarowski · 0 citations
Open access Sep 2026

Improvement and Application of Sampling Neural Network Using Virtual Space

Accurate nonlinear modeling underpins every layer of modern engineering. Sampling neural network (SNN) is a new fitting network, which has a simple structure, clear physical concept, concise algorithm, stable performance, and adopts a new error diffusion training method similar to the diffusion of neural stimuli in liv...

Ling-Yan Wu, Gang Cai · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.