Skip to content
Open access

Controllable Orthogonalization for Stabilizing Neural Network Training in Deep Reinforcement Learning

Oct 2026 · Informatics · 0 citations

Abstract

Deep reinforcement learning (DRL) relies on neural networks trained under temporally correlated experience, evolving replay distributions, and bootstrapped targets, conditions that can destabilize neural function approximation. This study investigates controllable orthogonalization as a network-level mechanism for improving training stability. Orthogonalization by Newton’s Iteration (ONI) is applied selectively to fully connected layers, with layer placement and Newton iteration count T controlling the location and strength of orthogonalization. We first analyze ONI placement and intensity in Deep Q-Network (DQN), and then evaluate the selected configurations on MinAtar and Atari. The same design principle is applied to the critic of Data-Regularized Q version 2 (DrQ-v2) on DeepMind Control Suite tasks. DQN-ONI improves aggregate normalized scores and learning trajectories across the evaluated discrete-control games. On the twelve-task continuous-control benchmark, Critic-ONI exhibits no learning failures. Learning-rate sweeps and penultimate-layer activation statistics further show a wider usable optimization range and more stable intermediate representations. These results show that selective, intensity-controlled orthogonalization can improve the training stability of DRL neural networks.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.