Skip to content

An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies

Jul 2026 · arXiv.org · Vol abs/2607.19771 · 0 citations · 8 references
Computer Science

TL;DR

A unified framework built on a single idealizing assumption -- exact scale invariance of the loss under weight rescaling, which holds approximately in normalization-heavy networks is proposed.

Abstract

Muon and related matrix-sign optimizers are increasingly used to pre-train large language models, but their effect on the internal geometry of individual weight matrices is not well understood. This preliminary report proposes a unified framework built on a single idealizing assumption -- exact scale invariance of the loss under weight rescaling, which holds approximately in normalization-heavy networks. Under this assumption, plain SGD carries a built-in 1/||W|| brake on its update size, whereas Muon's matrix-sign step removes that brake, so both the Frobenius and spectral norms drift outward faster (t^{1/2} versus t^{1/4}). We further observe that the spectral-norm perturbation has a non-negative second-order term. This implies that a lightweight"spectral cap"-- which projects out only the first-order growth of the single top singular direction from each update -- can control the output covariance W K_X W^T without freezing training: the weight keeps learning through non-top directions, top-direction rotation, and top switching. We relate this cap to the min-entropy (H-infinity) of the singular-value spectrum. We then study three systems trained with Muon: a nanoGPT feed-forward projection, a 64-expert mixture-of-experts router, and the query/key projections of a bf16 FlashAttention block. In each case the cap increases isotropy and, at the margins -- a router collapsing to a single expert, and the near-divergence of one attention head -- prevents a concrete failure, while leaving validation loss essentially unchanged. We emphasize that the scale-invariance assumption is strong and that these small-scale results are preliminary; comments are welcome.

View source

Similar papers

Jul 2026

Hyperball May Not Be a Free Lunch

An angular effective learning rate is derived that accounts for the parameter-update angle, parameter norm, and update norm, and shows that the conventional norm-based measure is a special case under parameter-update orthogonality.

Yihao Xiao, Jialong Sun, Zitian Gao et al. · 3 citations
Preprint Sep 2026

Musec: MomentUm SpEctral Clipping for Stable Muon-type Training

Muon has emerged as a highly effective optimizer for large language model training, often achieving superior convergence and performance compared with the widely adopted Adam and AdamW optimizers. Nevertheless, Muon is prone to training instability due to its spectral flattening, manifested by loss spikes and unbounded...

Zhuang-Hua Liu, Meng-Li Wang, Luo Luo · 0 citations
#machine learning Preprint Sep 2026

Sharp Rates and a One-Line Correction for Spectral Representation Learning

A self-supervised encoder is trained once, frozen, and reused through lightweight probes on tasks nobody named at training time; the practitioner's question is when the off-the-shelf features are good enough and when they need fixing. Canonical correlation analysis, HGR maximal correlation, and the population optimum o...

Di-Er Tang, Jing-Yee Tan, Guang-Yue Han · 0 citations
#artificial intelligence Preprint Sep 2026

Equivariance Breaks the Learning Rate

Equivariant networks are commonly trained with Adam, yet recent work reports that matrix-structured optimizers such as Muon can perform better on these architectures without explaining why. We identify one source of this difference inside equivariant linear layers. Each irrep block learns a channel-mixing matrix $W_l$...

Andrei Manolache, Mathias Niepert · 0 citations
Preprint Aug 2026

Machine-Learning Search for Lax Connections

We apply a machine learning framework to search for Lax connections in two-dimensional non-linear sigma models using local current data. For the $SU(2)$ principal chiral model and the symmetric coset $S^2 = SU(2)/U(1)$, the method successfully recovers the full spectral-parameter families without using the known spectr...

Osamu Fukushima, Tomohiro Shigemura, Ryosuke Suda et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.