This work introduces RODE, which gives the radial and directional components separate update rules and step sizes in the matrix Frobenius norm, and suggests that decoupling radial and directional dynamics offers a more effective and controllable approach to matrix optimization.
Abstract
Modern neural network training increasingly uses matrix-aware optimizers, yet their conditioned matrix step is typically added directly to the weight, jointly changing its norm and direction. This interaction matters because the current norm determines angular motion, while directional learning can drive norm growth and thereby alter later steps. We introduce RODE, which gives the radial and directional components separate update rules and step sizes. RODE explicitly updates the matrix Frobenius norm through a scalar radial rule, while its directional channel performs Newton--Schulz-conditioned updates in the tangent space. Controlled GPT-2 interventions show gains from both direct norm control and RODE's directional update. Across two language-modeling and two image-classification tasks, RODE outperforms both Muon variants in every direct comparison and ends with lower full-model norms. At 1.5B scale, using the learning rate transferred directly from the Qwen2-style LM sweep, RODE lowers loss from 4.145 to 3.346 and final global norm from 11964 to 2183 relative to Muon RMS, with fixed-radius RODE improving further. For Qwen3.5-9B full-parameter fine-tuning, all six optimizers use the same tuning budget and the same formal-training and evaluation settings; RODE outperforms both Muon variants on all four evaluation tasks and attains the highest mean on GSM8K and MATH-500. Thus, decoupling radial and directional dynamics offers a more effective and controllable approach to matrix optimization.
First-order gradient-based optimization is fundamental to training neural networks, yet standard adaptive and momentum-based methods primarily scale parameter updates using historical gradient magnitudes. As a result, they do not explicitly leverage the directional agreement between the current gradient and the accumul...
Rajeeb Thapa Chhetri, Saurab Thapa, Zhi-Xiong Chen et al.· International Conference on...· 0 citations
Low-Rank Adaptation (LoRA) is an effective approach for adapting large pretrained models by learning low-rank weight updates. In practice, the LoRA rank is used to control an adapter's parameter budget and representational capacity. We show that this view is incomplete: while the nominal rank determines the representat...
Zi-Han Zhu, Zhe-Hang Du, Xu-Yang Chen et al.· 0 citations
It is shown that a principled design space for LoRA initialization and curvature preconditioning should be treated as a tunable dimension rather than a fixed design decision, and a deployable, search-free variant, ULoRA-Auto, selects per-layer exponents from measured spectral statistics, approaches this upper bound at...
An angular effective learning rate is derived that accounts for the parameter-update angle, parameter norm, and update norm, and shows that the conventional norm-based measure is a special case under parameter-update orthogonality.
This paper proposes MOON (Multi-Objective OrthoNormalized Updates), which performs gradient manipulation under spectral--nuclear norm geometry and uses the orthonormalized manipulated gradient for parameter updates and empirical results show that MOON consistently improves both optimization efficiency and final multi-t...
Shiji Zhou, Kunlin Lyu, Lei Zhang et al.· 0 citations
Muon updates matrix-valued neural-network parameters by orthogonalizing a gradient-based momentum matrix. Its reliance on derivatives limits its use when gradients are unavailable or unreliable. We develop a derivative-free framework that constructs Muon-style updates from structured finite differences. Four variants a...
Peng-Cheng Xie· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.