A Novel Weight Initialization Scheme for Randomized Leaky Rectified Linear Units and a Comparative Analysis of Rectified Activation Functions for Image Classification
This study proposes a novel weight initialization scheme specifically designed for the randomized leaky rectified linear unit (RReLU) activation function, with the objective of preserving signal statistics during both forward and backward propagation.
Abstract
The choice of activation functions and weight initialization strategies plays a critical role in the trainability and performance of deep neural networks. In deep architectures, improper weight initialization can lead to vanishing or exploding signals and gradients, resulting in unstable training behavior. This study proposes a novel weight initialization scheme specifically designed for the randomized leaky rectified linear unit (RReLU) activation function, with the objective of preserving signal statistics during both forward and backward propagation. The proposed formulation accounts for the stochastic distribution of the slope in the negative input region of RReLU when deriving variance-preserving initialization conditions for both forward and backward propagation. The method is theoretically derived and experimentally evaluated on feedforward neural networks and convolutional neural networks (CNNs). Comparative evaluations against the widely used Xavier and He initialization schemes demonstrate that the proposed approach provides more stable optimization and facilitates successful convergence, particularly as network depth increases. In addition, rectified activation functions—including the rectified linear unit (ReLU), leaky rectified linear unit (LReLU), parametric rectified linear unit (PReLU), and RReLU—are systematically compared using their corresponding initialization strategies on the CIFAR-10, CIFAR-100, and ImageNet datasets. The experimental results show that ReLU remains the most computationally efficient activation function, whereas PReLU achieves the highest classification accuracy on the large-scale ImageNet dataset. On the relatively smaller CIFAR-10 and CIFAR-100 datasets, LReLU and RReLU exhibit competitive generalization performance. These findings highlight the importance of activation-specific initialization strategies for enabling fair comparisons and improving the performance of deep learning models.
Deep learning architectures comprise hierarchies in modern machine learning studies, and these are not only full of semantic depth; they are also structured systematically by a sequence of nonlinear transformations. In particular, in convolutional neural networks, the activation functions are the canonical nonlinear co...
Ali Raza, Akhtar Ali, Sami Ullah et al.· PLoS ONE· 0 citations
A simple and broadly applicable residual guided procedure that greedily constructs the hidden layer using a closed form residual decrease criterion and yields a progressive training process with a guaranteed monotonic decrease of the training objective.
Optimization plays a critical role in training deep neural networks, directly impacting convergence speed, model generalization, and stability. While existing methods such as stochastic gradient descent (SGD) and adaptive optimizers like Adam and AdamW have achieved significant success, they exhibit limitations in hand...
Ahmad Raza Khan, Sarab Almuhaideb· Electronics· 0 citations
Deep neural networks are commonly trained using adaptive optimization methods such as Adam because they converge quickly and perform well under stochastic training conditions. Despite these advantages, recent research has revealed several important drawbacks of Adam. In particular, the optimizer tends to converge towar...
Susilo Hariyanto, Siti Khabibah, Retno Putri Dwi Rahmawati et al.· WSEAS Transactions on System...· 0 citations
Activation functions play an important role in the performance of deep learning models by enabling them to learn complex and nonlinear relationships. However, widely used activation functions, such as ReLU, ELU, Swish, and GELU, have certain limitations. Issues such as exploding and vanishing gradients and dead neuro...
Özcan KüçükalI, M. H. Bozkurt, Esma Ulutas et al.· Concurrency and Computation· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.