Skip to content

Geodesics and Low Rank Behavior in the Deep Linear Network

· 0 citations · 29 references

TL;DR

A general system of ordinary differential equations describing geodesics in the DLN is derived and an investigation into using an entropic log-volume form related to the geometry on the full-rank manifold as an explicit regularizer for a simple class of energies is investigated.

View source

Similar papers

Sep 2026

The Geometric Whitney Problem and Approximations by Neural Networks on Manifolds

Abstract Why do neural networks overcome the curse of dimensionality? A common justification is that real-life high-dimensional data typically lie close to low-dimensional manifolds, and that neural networks can exploit this structure efficiently – overcoming the curse of dimensionality for their parameter counts. However, existing bounds depend on properties of the manifolds that cannot be read off from data alone. We close this gap. If a dataset locally looks like a low-dimensional linear space – a condition testable directly from data and derivable from the empirically supported manifold hypothesis under well-behaved conditions – then an approximating manifold M can be constructed. Neural networks can then approximate C1 functions uniformly on M, with parameter counts bounded purely in terms of computable properties of the data and overcoming the curse of dimensionality.

Jakob Konstantin Hecker · 0 citations
#machine learning Preprint Sep 2026

Symmetries and Singularities

Deep neural networks are highly over-parameterized, and different parameter values represent the same predictive function. This makes their effective complexity difficult to measure using only the number of parameters or the rank of the Hessian. Singular Learning Theory addresses this issue through the local learning coefficient (LLC), which characterizes the effective complexity of a model near a given solution. Existing methods for estimating the LLC often rely on posterior sampling, which can be computationally expensive for large neural networks. This makes accurate LLC estimation difficult at scale. In this work, we use known structures in the model to simplify the analysis and make LLC estimation more tractable. Specifically, we study the LLC of a graph attention model by exploiting symmetries in both the graph structure and the attention parameters. An analytic framework through a teacher--student setting, and explicit LLC estimates after considering the symmetry--induced degeneracies are developed.

Vishnu Varadarajan, Mihir More, Aritra Das et al. · 0 citations
Open access Aug 2026

High-order accurate inference on manifolds

We present a new framework for statistical inference on Riemannian manifolds that achieves high-order accuracy, addressing the challenges posed by non-Euclidean parameter spaces frequently encountered in modern data science. Our approach leverages a novel and computationally efficient procedure to reach higher-order asymptotic precision. In particular, we develop a bootstrap algorithm on Riemannian manifolds that is both computationally efficient and accurate for hypothesis testing and confidence region construction. Although locational hypothesis testing can be reformulated as a standard Euclidean problem, constructing high-order accurate confidence regions necessitates careful treatment of manifold geometry. To this end, we establish high-order asymptotics under an appropriate coordinate representation induced by a second-order retraction, thereby enabling precise expansions that incorporate curvature effects. We demonstrate the versatility of this framework across various manifold settings, including spheres, the Stiefel manifold, fixed-rank matrix manifolds, and rank-one tensor manifolds; for Euclidean submanifolds, we also introduce a class of projection-like coordinate charts with strong consistency properties. Finally, numerical studies confirm the practical merits of the proposed procedure.

Cheng-Zhu Huang, An-Ru R. Zhang · 0 citations
Preprint Aug 2026

Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws

It is shown that a symmetry fixes what variables are: a network layer is a sum over interchangeable units, so relabeling the units leaves it unchanged; given smoothness and the condition that a unit's gradient vanish at the origin, symmetry then enforces a universal leading form for the expansion about the near-zero weights present at the start of training.

Zi-Yin Liu, Yizhou Xu, Tomaso A. Poggio et al. · 0 citations
Preprint Aug 2026

Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks

This work expresses the known accessibility transition in an equivalent conic form, centered for compact convex targets at the statistical dimension of the polar cone, and introduces Random Mapping Networks (RaMaN), which instantiate the predicted latent dimension using structured Hadamard or seed-regenerated Gaussian maps.

Andrew Cheng, Ali Eslamian, Jie Cheng et al. · 0 citations
Jul 2026

Riemannian Deep Learning: Modules, Networks, and Geometries

Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific manifolds, rely on Euclidean approximations, or require costly and numerically fragile geometric operations. This thesis develops a unified framework for Riemannian deep learning from three complementary perspectives: reusable neural modules, manifold-specific network architectures, and the design of underlying geometries. It generalizes batch normalization from Euclidean spaces and individual manifolds to broad classes of Lie groups and gyrogroups, and extends multinomial logistic regression from Euclidean space to SPD manifolds and then to general Riemannian manifolds. It further develops neural networks for several important geometric representations, including an unconstrained model of hyperbolic space, Busemann-based hyperbolic learning, and full-rank correlation matrices. Finally, it introduces adaptive and computationally efficient Riemannian metrics on SPD manifolds, including learnable Log-Euclidean geometries and fast, stable Cholesky-based geometries. The proposed methods are supported by theoretical analysis and validated through numerical experiments and applications in vision, signal processing, graph learning, and genomics.

Ziheng Chen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.