Skip to content

Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods

Sep 2026 · 0 citations · 28 references
Computer Science

TL;DR

An Online Mirror Descent framework with adaptive proximal functions for matrix-valued parameters, providing a principled methodology for deriving matrix-aware adaptive optimization through online regret minimization and yields Row-wise Matrix AdaGrad and Column-wise Matrix AdaGrad as concrete instantiations with regret guarantees.

Abstract

Adaptive optimization methods such as AdaGrad and Adam are widely used in modern deep neural network training, but their adaptive scaling is primarily designed for vector-valued parameters and does not explicitly exploit matrix structure. Recent matrix-aware optimizers demonstrate the benefits of structured optimization, yet a general theoretical framework for deriving matrix-aware adaptivity comparable to that of AdaGrad remains lacking. In this work, we develop an Online Mirror Descent framework with adaptive proximal functions for matrix-valued parameters, providing a principled methodology for deriving matrix-aware adaptive optimization through online regret minimization. By introducing row-wise and column-wise matrix proximal functions, our framework explicitly reveals the trade-off governing adaptive scaling: increasing the scaling factors reduces the gradient-dependent dual norm term while increasing the cost of evolving the proximal geometry. In the row-wise setting, this trade-off becomes separable under diagonal parameterization, allowing the adaptive scaling for each row to be derived independently by minimizing its corresponding row-wise regret bound. The column-wise counterpart follows directly by applying the row-wise construction to the transposed matrix. This framework yields Row-wise Matrix AdaGrad and Column-wise Matrix AdaGrad as concrete instantiations, with regret guarantees that are strictly tighter than those of entry-wise AdaGrad under row-sparse or column-sparse gradient structures. Experiments on matrix factorization and stacked deep MLP training further demonstrate the benefits of matrix-aware adaptive scaling, yielding improved optimization performance in both settings and enhanced optimization stability and trainability at larger learning rates and greater network depths in the latter.

View source

Similar papers

Preprint Aug 2026

Accelerated Learning of High Dimensional Functions with a Tensor-Featured Training Network

This work presents a method to accelerate the optimization of learning high dimensional functions using deep neural network (DNN), and studies the effect of adding features which distill pretrained DNN into TNs using a discretize and decompose strategy.

Karl Pierce, Y. Khoo, Haizhao Yang · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond the Matrix Sign: Quadratic Spectral Descent

Muon emerges as a strong competitor of the AdamW for LLM pretraining, because the matrix-wise update it employs can potentially incur smaller second-order penalty than the once dominating AdamW, which performs coordinate-wise update. However, the spectral flattening procedure in Muon is quite debatable since it discard...

Qiao-Zhe Zhang, Jun Sun, Ying-Zhuang Liu · 0 citations
Preprint Sep 2026

Gradient-Free Optimization for Matrix functions

An alternative to the standard random gradient estimator is introduced, allowing for the projection step of spectral descent to be done at no extra cost and it is shown that by exploiting this low-rank property one obtains much faster convergence to good approximate solutions.

S. Allen, Cash Cherry, Aidan Eck et al. · 0 citations
Preprint Aug 2026

On the Principles Behind Neural Network Optimizers

This thesis develops a principled grounding for Adam and motivates new designs, and reveals new local structures in matrix-based nonconvex problems, and helps understand and improve recent NN optimizers, such as Muon.

Yu-Shun Zhang · 0 citations
Preprint Aug 2026

RODE: A Radial-Orthogonal Decoupled Engine for Optimization

This work introduces RODE, which gives the radial and directional components separate update rules and step sizes in the matrix Frobenius norm, and suggests that decoupling radial and directional dynamics offers a more effective and controllable approach to matrix optimization.

Guo-Xiang Xu, Bince Qu, Qi Sun et al. · 0 citations
Preprint Aug 2026

The Loss Does Not See the Basis, but Adam Does

Gradient descent on a factored model $W = UV^\top$ is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization, is not. We trace the difference to the gauge symmetry of the loss, its invariance under $(U, V) \mapsto (UQ, VQ)$. Gradient flow's low-rank mechanism is available t...

D. Singh · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.