Skip to content
Preprint

The Loss Floor of Denoising Score Matching: Fisher Geometry from Schr\"odinger Bridges

Aug 2026 · 0 citations · 83 references
Computer Science Physics

TL;DR

Denoising score matching trains diffusion models by regressing onto a conditional score, although generation ultimately requires the marginal score, and derives the result from a Schr"odinger bridge variational principle, in which the ideal objective arises as excess path-space relative entropy.

Abstract

Denoising score matching trains diffusion models by regressing onto a conditional score, although generation ultimately requires the marginal score. The two objectives share the same population minimizer, but the conditional target remains random at fixed noisy state and introduces an irreducible excess in the training loss. We isolate this excess and show that, for a general corruption kernel under mild regularity assumptions, it is exactly the trace of the Fisher--Rao metric of the conditional endpoint family, integrated along the diffusion trajectory. This gives an exact conditional-variance decomposition of the denoising objective and identifies the information geometry observed in diffusion latent spaces as an intrinsic component of the training loss. We derive the result from a Schr"odinger bridge variational principle, in which the ideal objective arises as excess path-space relative entropy. For corruption diffusions, the Fisher term is proportional to the rate at which the noisy state loses mutual information about the clean data, separating the loss floor into an information flow determined by the data and a weight determined by the corruption schedule and objective. In the Gaussian case, this yields a closed form for the floor, recovers reparametrization invariance of the continuous-time objective, and relates its high-SNR divergence to the information dimension of the data. Finally, we show that raw losses obtained with different noise ranges or weightings need not rank models consistently because they contain different additive floors, and contrast the second-order geometry seen by training with the third-order conditional statistics entering numerical sampling error.

View source

Similar papers

#machine learning Preprint Sep 2026

Noise2Noise Revisited: Training Pair Distributions Dominate Loss Choice in Self-Supervised Denoising

Noise2Noise (N2N) trains denoisers on pairs of independently corrupted observations, eliminating clean references. We stress-test two natural conjectures about why the L1 loss outperforms L2 here. First, the hypothesis that the L1 loss confers robustness via parameter sparsity confuses the loss with Lasso regularizatio...

Ding-Yan Shang, Zhen-Yu Xu, You Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Sharp Rates and a One-Line Correction for Spectral Representation Learning

A self-supervised encoder is trained once, frozen, and reused through lightweight probes on tasks nobody named at training time; the practitioner's question is when the off-the-shelf features are good enough and when they need fixing. Canonical correlation analysis, HGR maximal correlation, and the population optimum o...

Di-Er Tang, Jing-Yee Tan, Guang-Yue Han · 0 citations
Preprint Aug 2026

Drift Variation Autoencoder: Unifying Generation and Representation Learning through Conditional Posterior Flow Matching

Stochastic masking, cropping, or modality removal makes deterministic reconstruction an incomplete target: one observation can admit many clean completions. This work takes the corresponding posterior $P(X\mid C)$ as the common statistical object for conditional generation and generatively sufficient representation lea...

Jiarui Cao · 0 citations
#machine learning Preprint Sep 2026

Why Learning Rediscovers the Closed-Form Diagonal Regularizer

We identify a diagonal saturation principle in modal inverse problems: when truncation noise is isotropic, the Bayes-optimal Tikhonov shape is a closed-form power law Gamma_k proportional to lambda_k^|s| set by the prior alone, independent of the domain. Berry's random-wave conjecture decorrelates the truncation noise...

Jeahn Han, Pyojin Kim · 0 citations
#artificial intelligence Preprint Sep 2026

Let It Go or Learn to Self-Correct: Continuous Diffusion for Constrained Discrete Tasks

Denoising Diffusion Probabilistic Models (DDPMs) generate samples by starting from noise and repeatedly denoising while keeping each update close to the current noisy state. This behavior is effective in many continuous domains, but its role is less clear for globally constrained discrete tasks, such as Sudoku, graph c...

M. Drozdova, Stéphane Liem Nguyen, François Fleuret · 0 citations
#machine learning Preprint Sep 2026

Explaining f-Divergence-Based Regularization via Local Curvature and Sharpness-Aware Minimization

Divergence-based regularization and Sharpness-Aware Minimization (SAM) are two prominent approaches for improving generalization in deep learning, both motivated by robustness to perturbations. However, their relationship has remained largely unexplored. Building on classical second-order expansions of $f$-divergences,...

Nour Jamoussi, Marios Kountouris · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.