Skip to content
Preprint

Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks

Aug 2026 · 0 citations
Computer Science

TL;DR

This work expresses the known accessibility transition in an equivalent conic form, centered for compact convex targets at the statistical dimension of the polar cone, and introduces Random Mapping Networks (RaMaN), which instantiate the predicted latent dimension using structured Hadamard or seed-regenerated Gaussian maps.

Abstract

Neural networks can often be trained or fine-tuned through random low-dimensional reparameterization, where a small latent vector is mapped into a full parameter update by a frozen random map. This raises a practical question: how large must the latent search space be to reach a low-loss region? We first express the known accessibility transition in an equivalent conic form, centered for compact convex targets at the statistical dimension of the polar cone. Our main theoretical contribution is an orientation-resolved quadratic master formula that predicts the random-slice residual from both the curvature spectrum and the reference-to-solution displacement profile. It yields a self-consistent isotropic-orientation predictor and, in a conservative radius-only specialization, recovers the earlier Gaussian-width quadratic bound. Building on this analysis, we introduce Random Mapping Networks (RaMaN), which instantiate the predicted latent dimension using structured Hadamard or seed-regenerated Gaussian maps. These constructions avoid the O(dP) storage of dense random maps and reduce optimizer-state memory from O(P) to O(d). We also develop matrix-free curvature approximations and sweep-free dimension selection. Across controlled quadratic and neural-curvature experiments, the orientation-resolved predictor closely tracks measured transition locations and outperforms orientation-agnostic approximations when displacement direction matters. End-to-end experiments further show sharp, protocol-dependent training transitions across image and language models.

View source

Similar papers

Geodesics and Low Rank Behavior in the Deep Linear Network

A general system of ordinary differential equations describing geodesics in the DLN is derived and an investigation into using an entropic log-volume form related to the geometry on the full-rank manifold as an explicit regularizer for a simple class of energies is investigated.

Alan Chen, Govind Menon · 0 citations
Jul 2026

Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries

Delayed generalization, or grokking, remains poorly understood despite extensive empirical study. We identify an exactly solvable late-time relaxation mechanism for grokking in linear models trained with full-batch heavy-ball optimization and weight decay, together with a locally quadratic extension to nonlinear neural networks. Our analysis reveals a distinguished population-active component of the empirical null space, which we call the grokking subspace. Along this subspace, the training predictions remain unchanged, leaving weight decay as the sole restoring force and giving rise to a slow dissipative relaxation governed by an exact discrete-time and continuous-time law. We show that only this subspace contributes to the slow asymptotic decay of the population risk and derive explicit iteration-scale predictions for the grokking time, recovering the familiar $(1-\beta)/(\eta\lambda)$ scaling in the weak-regularization regime. The theory further predicts distinct effects of optimizer choice, distinguishing coupled $L_2$ regularization from decoupled weight decay, and yields causal predictions for interventions that modify the grokking component. We verify all theoretical identities without fitted parameters in a synthetic model where every subspace and relaxation rate is computable in closed form. We further observe genuine delayed generalization in modular addition, where the measured delay follows the predicted scaling and the late-time relaxation agrees closely with the theoretical clock.

Taeyoung Kim · 1 citation
Preprint Sep 2026

Machine-Learned Dynamical Representations for Accelerated RiteWeight Convergence

The increasing use of generative models has made ensembles of short molecular dynamics trajectories increasingly common, creating a growing need for methods that can recover physically meaningful steady-state populations and kinetics from improperly weighted conformational ensembles. Randomized Iterative Trajectory Reweighting (RiteWeight) addresses this problem through repeated random clustering and iterative reweighting, without requiring the fixed Markovian discretization used in conventional Markov state models (MSM). However, the choice of reduced feature space in which RiteWeight performs random clustering has not been systematically investigated. Here, we compare two machine-learned representations, DeepTICA and SPIB-VAE, with linear TICA for recovering steady-state observables from flawed distributions. DeepTICA learns nonlinear coordinates by targeting slow transfer-operator eigenmodes, whereas SPIB-VAE compresses configurations into a low-dimensional latent space while retaining information predictive of future metastable states. DeepTICA provided a comparatively robust RiteWeight representation under limited hyperparameter exploration, whereas SPIB-VAE benefited more strongly from broader optimization. Moreover, a kinetic score computed from a coarse MSM at a resolution comparable to that used for RiteWeight random clustering provided a useful criterion for efficiently selecting reduced representations and their hyperparameters for RiteWeight.

Sagar Kania · 0 citations
Preprint Aug 2026

TESLA: Taylor Expansion of Sinusoidal Learnable Activations

TESLA, an activation defined as a learnable combination of sine and cosine terms, enabling explicit control over polynomial degree and selective amplification of high-order components is proposed, indicating that activation-level degree control transfers to more general vision workloads.

Daehwa Ko, Jae-Hwan Kim, Seunghyun Ham et al. · 0 citations
#machine learning Preprint Aug 2026

Residual-Guided Randomized Neural Networks

A simple and broadly applicable residual guided procedure that greedily constructs the hidden layer using a closed form residual decrease criterion and yields a progressive training process with a guaranteed monotonic decrease of the training objective.

M. Akhtar, M. Tanveer, Mohd. Arshad · 0 citations
#machine learning Preprint Sep 2026

Structured Features Overfit Where Random Features Grok

Xu, Vardi and Safran (ICML 2026) prove that over-parameterized ridge regression over an unstructured random Gaussian feature map groks, with the delay between memorization and generalization growing as $1/\lambda$ in the weight decay. We show that on a structured feature map the same delay does not appear. For a band-limited Fourier feature map over $\mathbb{Z}_p^2$ carrying a single-character target that lies inside the expressible class, enlarging the band at fixed positive weight decay drives peak held-out accuracy monotonically from $1.00$ to $0.07$, with no memorize-then-generalize regime anywhere along the sweep. The degradation is not an interpolation effect. It sets in at capacity ratio $q/n = 0.638$, far below the interpolation threshold, on separate grounds from the exact null space that appears above it. What does have a sharp boundary is the active support. Holding the nominal dimension fixed and masking the band back to $1089$ active modes restores held-out accuracy of $1.000$ with zero variance across seeds, while the full $4225$-mode band collapses to $0.185$. The number of active modes acts through the teacher-weighted spectrum of the empirical Gram matrix and not through the capacity ratio, which makes this a statement about feature geometry and not a restatement of double descent.

Chon-Fai Kam, M. Bessafi, Frédéric Cadet · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.