Skip to content

Conservation Laws for Diffusion Models

Jul 2026 · arXiv.org · Vol abs/2607.10067 · 0 citations · 33 references
Computer Science Mathematics

Abstract

While autoregressive models optimize the exact data likelihood via the chain rule, diffusion models are typically trained with denoising objectives. We develop conservation laws based on generalized extrinsic information transfer (GEXIT) functions for a broad class of memoryless noise processes, showing that the data--model cross-entropy (CE) can be characterized exactly as an integral of local information-theoretic derivatives along the noise path. This yields a unified characterization of the likelihood for discrete and continuous diffusion, with the Gaussian case reducing to the well-known mutual information--minimum mean-square error (I-MMSE) relationship. An immediate implication is a locality property: one can compute the information-theoretic derivatives using only the marginal posteriors along the noise path. As a result, training reduces to learning the marginal posteriors by minimizing the negative log-likelihood. While the conservation law implies that the entropy does not depend on the noise path, finite-capacity denoisers approximate the posteriors with varying accuracy across noise types, leading to differences in performance. We validate these predictions on synthetic Markov sources and standard benchmarks, including text8 and CIFAR-10.

View source

Similar papers

Preprint Aug 2026

Posterior Information Dynamics of Diffusion Models for Linear Inverse Problems

Diffusion models are widely used as priors for linear inverse problems, yet endpoint quality does not reveal when measurement information enters reverse denoising or how it is allocated across signal directions. We study this process through the smoothed likelihood force, the difference between exact posterior and prio...

Xiangming Meng · 1 citation
Aug 2026

Improved denoising diffusion probabilistic models with efficient non-diagonal covariance modeling

The Kronecker-DCT (K-DCT) model uses a Kronecker-factored decomposition of inter-color covariances and spatial covariances modeled in the frequency domain using the Discrete Cosine Transform (DCT), resulting in negligible computational and memory overhead in each denoising step.

Rui Xia, Ayan Das, A. Artemev et al. · 0 citations
Preprint Aug 2026

Forward-Evolution Error Analysis and Adaptive Design for Matrix-Valued Diffusion Models

This analysis transfers reverse-time discretization errors to the forward corruption law and treats two numerical schemes within a common framework, which identifies a local error, accumulates it through the forward evolution, and inserts the result into a common KL decomposition.

T. Pang, Zuowei Shen, Rui-Tong Zhang · 1 citation
Preprint Aug 2026

Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime

This work develops a generative counterpart to the theory of benign overfitting and algorithmic regularization for overparameterized neural networks in the supervised lazy-training regime by studying denoising score matching in a vector-valued reproducing kernel Hilbert space with an inner-product kernel.

Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian et al. · 1 citation · ⚡1
Preprint Aug 2026

Diffusion Quasi-Monte Carlo

This work constructs a cube-to-target map by composing a Gaussian base transformation (the component-wise inverse Gaussian CDF) with an Euler-discretized probability flow ODE, and establishes conditions for diffusion probability-flow transport under mild bounded-derivative assumptions on the learned vector field.

Jian-Long Chen, Yi-Feng Yu · 0 citations
#machine learning Preprint Aug 2026

Exact Global MCMC with Denoising Diffusion

This work shows that diffusion models learned with standard denoising loss can provide effective global MCMC proposals for complex high-dimensional target densities and offers preliminary evidence that the established scaling behavior of standard diffusion training transfers directly to exact sampling from high-dimensi...

Mitch Hill · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.