Skip to content

Phasor Attention: Mean Root Square Normalization for Phase Manifold Preservation

Jul 2026 · arXiv.org · Vol abs/2607.17822 · 0 citations · 12 references
Computer Science

TL;DR

Under extreme hyperparameter settings where standard normalizations suffer from gradient divergence, MRSNorm successfully prevents numerical explosion and secures stable optimization trajectories, proposing a fundamental paradigm shift toward phasor-based deep representation learning.

Abstract

While Root Mean Square Normalization has become the de facto standard for accelerating modern sequence models, its reliance on the quadratic accumulation of independent scalars ($\sum x^2$) inherently triggers outlier-induced numerical instability, gradient starvation, and anisotropic phase distortion. We introduce Mean Root Square Normalization (MRSNorm). By structurally pairing channels into 2D phasors, MRSNorm mathematically inverts the traditional scaling paradigm: it computes the localized $L_2$ magnitudes (Root Square) before aggregating them via a global $L_1$ average (Mean). This operational inversion strictly constrains activations to a phasor manifold, preserving conformal invariance. By sharing a single affine weight across phasor components, MRSNorm halves the total number of learnable parameters, proving that unconstrained spatial scaling in standard norms is a harmful redundancy. We analytically demonstrate that this geometric constraint yields a built-in, trigonometric gradient clipper governed by the Pythagorean identity, unconditionally equalizing the local gradient norm to ensure Gradient Homogeneity. Empirical evaluations on a ResNet with CIFAR-100 show that despite halved parameters, MRSNorm provides critical structural stability under rigorous stress tests. Under extreme hyperparameter settings where standard normalizations suffer from gradient divergence, MRSNorm successfully prevents numerical explosion and secures stable optimization trajectories. Our findings propose a fundamental paradigm shift toward phasor-based deep representation learning. The implementation of MRSNorm is available at Appendix C.

View source

Similar papers

Preprint Aug 2026

Orthogonal Polynomial Approximation for Matrix Log Normalization in Global Covariance Pooling

Global Covariance Pooling (GCP) improves deep networks by capturing second-order feature statistics, and is especially effective for fine-grained recognition. Because covariance matrices live on the Symmetric Positive Definite (SPD) manifold, a normalization step is required before the Euclidean classifier. The faithfu...

Md Rifat Ur Rahman, Md Raihan Khan, Md Sakib Hossain Shovon et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers

Smaller activation outliers do not necessarily imply better low-bit quantization: their alignment with the quantizer matters. We introduce PrismQuant, a quantizer-aware rotation framework that aligns the leading activation eigenspace with the constant group subspace of asymmetric grouped INT4. The affine offsets repres...

Yan-Long Chen, Yi-Ning Chen, Song Zhang et al. · 0 citations
Preprint Aug 2026

Iterative Erasure Count Is Not an Affine-Invariant Concept Dimension

How many directions does a neural representation use to encode a concept? A common answer repeatedly erases probe directions and reports the stopping count or cumulative removed rank. We show that both quantities can change under an information-preserving invertible reparameterization, so neither is intrinsically a con...

Tingan Jin, Shu-Hang Dong, Hao-Song Li et al. · 0 citations
Open access Sep 2026

Affine-Intensity Invariance and Spatial Non-Identifiability in Spectral Channel Gates

Spectral channel attention compresses a spatial feature map into a weight, but invariance of that weight need not preserve evidence about spatial arrangement. This study asks whether centered spectral concentration can provide exact affine-intensity invariance and identify spatially altered inputs. A normalized fourth...

Ashaq Ali · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond the Matrix Sign: Quadratic Spectral Descent

Muon emerges as a strong competitor of the AdamW for LLM pretraining, because the matrix-wise update it employs can potentially incur smaller second-order penalty than the once dominating AdamW, which performs coordinate-wise update. However, the spectral flattening procedure in Muon is quite debatable since it discard...

Qiao-Zhe Zhang, Jun Sun, Ying-Zhuang Liu · 0 citations
Preprint Sep 2026

SubRot: Signed Gradient Subspace Calibration for VLM Rotation Quantization

Post-training quantization reduces the deployment cost of vision-language models (VLMs), but preserving multimodal capabilities at low bit widths remains challenging. Existing methods rely on modality- or token-level gradient statistics, which are susceptible to cross-sample variations in visual-to-textual token ratios...

Zhen-Hao Shang, Hai-Zhao Jing, Hao-Kui Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.