Controllable Orthogonalization for Stabilizing Neural Network Training in Deep Reinforcement Learning
Deep reinforcement learning (DRL) relies on neural networks trained under temporally correlated experience, evolving replay distributions, and bootstrapped targets, conditions that can destabilize neural function approximation. This study investigates controllable orthogonalization as a network-level mechanism for impr...