Skip to content
Preprint

Statistical Mechanics of Learning on Product Wasserstein Manifolds

Aug 2026 · 0 citations · 26 references
Computer Science Mathematics Biology

TL;DR

This approach reframes structural constraints as geometric priors and suggests a route for incorporating biological, spectral, or hardware-derived distributional information into both learning systems, viz., classical and quantum learning.

Abstract

Normally the statistical mechanics of learning treats constraints on weight distributions as restrictions that shrink the space of possible solutions. Therefore, it reduces model capacity. In this paper we would like to take a contrary approach, which, however, is based on the earlier work on distribution-constrained perceptrons. Rather than treating a prescribed weight distribution as a mere restriction, we propose that it defines the intrinsic geometry upon which learning naturally unfolds. We formulate both deep neural networks and variational quantum circuits as gradient flows on a product of Wasserstein manifolds -- one classical Wasserstein space for each layer and one quantum Wasserstein space for the circuit parameters. Within this geometry, the capacity reduction, which was previously associated with distributional constraints, appears as the metric structure of the constraint manifold itself. We develop a hierarchical mean-field description for deep networks, extend the framework to the quantum setting using the quantum Wasserstein distance of order 1, and introduce two such practical algorithms, Hierarchical DisCo-SGD and Quantum DisCo, that follow approximate geodesics on the manifold of the product itself. Experiments on teacher-student problems, standard image classification tasks, and small variational quantum classifiers show that respecting these distributional geometries improves generalization, stabilizes training, and reduces the severity of barren plateaus compared with unconstrained and purely norm-based baselines. This approach firstly reframes structural constraints as geometric priors and suggests a route for incorporating biological, spectral, or hardware-derived distributional information into both learning systems, viz., classical and quantum learning.

View source

Similar papers

#machine learning Preprint Sep 2026

When Riemann flows with Wasserstein: Generative Modeling of Probability Distributions on Manifolds

Many scientific datasets, such as molecular conformational ensembles or single-cell tissue measurements, are naturally modeled as meta-distributions: distributions over probability measures on non-Euclidean domains. Existing generative methods largely assume Euclidean geometry and fail to capture this structure. We int...

D. Haviv, E. De Brouwer, Rishabh Anand et al. · 0 citations
Nov 2026

Statistical inference for Bures-Wasserstein flows

A statistical framework for conducting inference on collections of time-varying covariance operators (covariance flows) over a general, possibly infinite dimensional, Hilbert space is developed, and asymptotic theory is fully developed based on interpretable and verifiable assumptions.

Leonardo V. Santoro, Victor M. Panaretos · 0 citations
#machine learning Preprint Sep 2026

RW-Flow: One-Step Generation on Compact Manifolds via Wasserstein Gradient Flows

Manifold-valued data, and consequently the distributions they induce, are prevalent across many domains, ranging from the locations of geospatial events, such as earthquakes, to biomolecular torsion angles that encode information about three-dimensional structure. While diffusion and flow-based generative models have b...

Ualibyek Nurgulan, Seungwoo Yoo, Prin Phunyaphibarn et al. · 0 citations
Preprint Sep 2026

Discrete Gromov-Wasserstein Duality: Algorithms and Isomorphism Testing

The Gromov-Wasserstein (GW) distance provides a principled framework for aligning metric measure (mm) spaces based solely on their intrinsic structure. Its ability to identify isomorphic representations of distributions across spaces renders it valuable for comparing data where equality up to isomorphism occurs natural...

Gabriel Rioux, Joanna Marks, Riccardo Passeggeri et al. · 1 citation
#machine learning Review Aug 2026

Rotational Equivariance in Machine Learning: A Comprehensive Tutorial

This tutorial provides a comprehensive introduction to rotational equivariance, starting from the physical and geometric intuition behind coordinate independence and building up the necessary machinery from geometric deep learning, group theory, and representation theory.

Peter Lippmann, Fred A. Hamprecht · 0 citations
Preprint Aug 2026

Gromov-Wasserstein Quantization and Clustering: Structure, Rates, and Algorithms

Numerical experiments show that GW quantization opens up many modeling possibilities beyond normal clustering methods and that the introduced algorithm leads to useful numerical solutions with approximation quality often in line with theoretically optimal rates.

F. Beier, S. Eckstein · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.