Skip to content
Open access

A Novel Weight Initialization Scheme for Randomized Leaky Rectified Linear Units and a Comparative Analysis of Rectified Activation Functions for Image Classification

Sep 2026 · Electronics · Vol 15, pp. 4325 · 0 citations · 33 references

TL;DR

This study proposes a novel weight initialization scheme specifically designed for the randomized leaky rectified linear unit (RReLU) activation function, with the objective of preserving signal statistics during both forward and backward propagation.

Abstract

The choice of activation functions and weight initialization strategies plays a critical role in the trainability and performance of deep neural networks. In deep architectures, improper weight initialization can lead to vanishing or exploding signals and gradients, resulting in unstable training behavior. This study proposes a novel weight initialization scheme specifically designed for the randomized leaky rectified linear unit (RReLU) activation function, with the objective of preserving signal statistics during both forward and backward propagation. The proposed formulation accounts for the stochastic distribution of the slope in the negative input region of RReLU when deriving variance-preserving initialization conditions for both forward and backward propagation. The method is theoretically derived and experimentally evaluated on feedforward neural networks and convolutional neural networks (CNNs). Comparative evaluations against the widely used Xavier and He initialization schemes demonstrate that the proposed approach provides more stable optimization and facilitates successful convergence, particularly as network depth increases. In addition, rectified activation functions—including the rectified linear unit (ReLU), leaky rectified linear unit (LReLU), parametric rectified linear unit (PReLU), and RReLU—are systematically compared using their corresponding initialization strategies on the CIFAR-10, CIFAR-100, and ImageNet datasets. The experimental results show that ReLU remains the most computationally efficient activation function, whereas PReLU achieves the highest classification accuracy on the large-scale ImageNet dataset. On the relatively smaller CIFAR-10 and CIFAR-100 datasets, LReLU and RReLU exhibit competitive generalization performance. These findings highlight the importance of activation-specific initialization strategies for enabling fair comparisons and improving the performance of deep learning models.

Read PDF

Similar papers

Open access Aug 2026

Development of adaptive activation functions with curvature and range modulation for abstract image feature learning in complex CNNs for cross-domain applications

Deep learning architectures comprise hierarchies in modern machine learning studies, and these are not only full of semantic depth; they are also structured systematically by a sequence of nonlinear transformations. In particular, in convolutional neural networks, the activation functions are the canonical nonlinear co...

Ali Raza, Akhtar Ali, Sami Ullah et al. · 0 citations
#machine learning Preprint Aug 2026

Residual-Guided Randomized Neural Networks

A simple and broadly applicable residual guided procedure that greedily constructs the hidden layer using a closed form residual decrease criterion and yields a progressive training process with a guaranteed monotonic decrease of the training objective.

M. Akhtar, M. Tanveer, Mohd. Arshad · 0 citations
Open access Aug 2026

Quartz Optimizer: Robust Gradient Shaping and Bounded Adaptive Steps for Stable Deep Learning Training

Optimization plays a critical role in training deep neural networks, directly impacting convergence speed, model generalization, and stability. While existing methods such as stochastic gradient descent (SGD) and adaptive optimizers like Adam and AdamW have achieved significant success, they exhibit limitations in hand...

Ahmad Raza Khan, Sarab Almuhaideb · 0 citations
Open access Sep 2026

NeuroFuzzyAdam: A Fuzzy-Enhanced Adaptive Optimization Algorithm for Deep Learning

Deep neural networks are commonly trained using adaptive optimization methods such as Adam because they converge quickly and perform well under stochastic training conditions. Despite these advantages, recent research has revealed several important drawbacks of Adam. In particular, the optimizer tends to converge towar...

Susilo Hariyanto, Siti Khabibah, Retno Putri Dwi Rahmawati et al. · 0 citations
Open access Sep 2026

A Novel Activation Function for Enhancing Deep Neural Network Stability and Performance

Activation functions play an important role in the performance of deep learning models by enabling them to learn complex and nonlinear relationships. However, widely used activation functions, such as ReLU, ELU, Swish, and GELU, have certain limitations. Issues such as exploding and vanishing gradients and dead neuro...

Özcan KüçükalI, M. H. Bozkurt, Esma Ulutas et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.