This analysis transfers reverse-time discretization errors to the forward corruption law and treats two numerical schemes within a common framework, which identifies a local error, accumulates it through the forward evolution, and inserts the result into a common KL decomposition.
Abstract
Diffusion models learn to reverse a predefined corruption process, but sampling still requires a costly time discretization and depends on the chosen noise schedule. We study these two issues for variance-preserving diffusions with matrix-valued schedules. Our analysis transfers reverse-time discretization errors to the forward corruption law and treats two numerical schemes within a common framework. The first freezes the score and yields, through a matrix-sensitive local comparison and forward information dissipation, an ambient-dimensional step complexity with leading factor $d/\varepsilon^2$ for KL accuracy $\varepsilon^2$. The second keeps the known Gaussian drift exact and freezes the posterior mean. For data of metric-entropy dimension $k$, a forward Markov identity, an anisotropic covering estimate, and Stieltjes integration by parts give the corresponding factor $k\log k/\varepsilon^2$. In both cases, the proof identifies a local error, accumulates it through the forward evolution, and inserts the result into a common KL decomposition. The local errors further provide directional criteria for matrix schedules and an asymptotically optimal square-root adaptive grid. A high-dimensional Gaussian-mixture experiment illustrates the resulting schedule and grid improvements.
Diffusion models are widely used as priors for linear inverse problems, yet endpoint quality does not reveal when measurement information enters reverse denoising or how it is allocated across signal directions. We study this process through the smoothed likelihood force, the difference between exact posterior and prior scores at each noise level. For a fixed measurement, its expected squared norm gives both posterior--prior relative-entropy dissipation and reverse-path relative-entropy growth. Averaging over measurements yields an information--minimum mean-square error (I-MMSE) identity linking information gain to denoising-error reduction. Under finite second moments, the force energy and its ratio to prior-score energy decay quadratically in the noising kernel's signal coefficient at high noise. Solvable models show that conditioning removes class separation already explained by the measurement, reduces a uniform index entropy over \(n\) empirical samples from \(\log n\) to \(H(I\mid r)\), and makes assimilation depend on operator--prior alignment even for identical singular values. Experiments in models with tractable posteriors evaluate these predictions. In a separate illustration with a frozen FFHQ model, masks sharing the same spectrum yield different prior-normalized null-space trajectory statistics.
This work constructs a cube-to-target map by composing a Gaussian base transformation (the component-wise inverse Gaussian CDF) with an Euler-discretized probability flow ODE, and establishes conditions for diffusion probability-flow transport under mild bounded-derivative assumptions on the learned vector field.
We introduce a unified framework for the error analysis of generative models based on the entropy production rate of the forward-reverse diffusion process pair. For a pair of continuity equation flows, the rate admits a closed velocity form identity whose time integral decomposes the terminal Kullback--Leibler (KL) divergence into the sum of an initialization error, a score approximation error, and a time-discretization error. By analyzing the entropy production at the level of marginal distributions, rather than in path space, our framework yields a sharp convergence rate of $\mathcal{O}(h^2)$ for the Euler-Maruyama sampler, where $h$ is the step size. This improves upon the $\mathcal{O}(h)$ rates typically obtained from Girsanov's path-space analyses. Furthermore, our framework unifies the analysis of score-based SDEs, probability-flow ODEs, and stochastic interpolants by varying diffusion coefficients within a single inequality, revealing the trade-off between deterministic and stochastic sampling. Numerical experiments confirm the predicted scaling with step size and terminal time.
This work considers a first-order sampler based on the leave-one-out denoiser for uniform and remasking processes whose coordinate updates can be performed in parallel, and derives an exact information-theoretic representation of the discretization error in terms of the mutual information between different coordinates of the forward process at different times.
Two central challenges in diffusion-based sampling are the theoretical one of understanding their remarkable effectiveness even in high-dimensional settings, and the practical one of designing algorithms with certified performance guarantees. We show that these questions are intimately connected via the \emph{denoising growth complexity} ($\mathsf{DGC}$). It is a geometric measure defined by a log-time weighted integral of the derivative of the denoising mean-squared error along the Gaussian heat flow. We show how the $\mathsf{DGC}$ increments lead to a simple and explicit bound on the KL error of an Euler scheme applied to the stochastic innovations representation. The bound is local along the path: each step is controlled by the corresponding $\mathsf{DGC}$ increment and its relative stepsize. This structure allows us to derive KL sampling guarantees for optimized stepsize schedules, both in a simpler single-block setting and in a more refined $K$-block setting. The $\mathsf{DGC}$ function has a natural martingale structure, which we exploit to develop fully data-certified versions of these algorithms. It also admits information-theoretic upper bounds in terms of covariance, rate distortion, metric entropy, and the Poincar'e constant, thereby recovering and sharpening a range of existing diffusion-sampling guarantees, as well as giving new results. In log heat-time, the fine partition limit is governed by an integral involving the square root of the $\mathsf{DGC}$ density, whereas a single-block schedule depends on its ordinary integral. This comparison precisely characterizes when adaptation to data geometry yields substantial computational gains, including logarithmic-to-constant separations for simple Gaussian mixture models.
We study propagation of chaos for decoupled mean-field forward-backward stochastic differential equations whose generators depend on the empirical laws of the forward states, backward values and diagonal martingale integrands. Under monotonicity and Lipschitz assumptions, synchronous coupling gives quantitative estimates, including an $m$-particle squared Wasserstein bound of order $m/n$ for a system of $n$ particles interacting through finitely many statistics. In the Markovian setting, assuming a sufficiently regular classical decoupling field, we obtain two sharp refinements. For constant invertible diffusion and first-order cancellation of the field's measure dependence along the limiting law flow, the squared Wasserstein error is of order $m^2/n^2$, on continuous-path space for the values and on $L^2$ for the diagonal integrands. Without imposing this cancellation, smooth weak errors have order $n^{-1}$ for every fixed marginal, allowing variable and possibly degenerate diffusion. The weak estimate is uniform on a fixed time interval for the values and integrated in time for the integrands. The argument compares the interacting BSDE with an empirical evaluation of the decoupling field, retaining the full martingale representation and controlling feedback through both backward laws. It yields a joint-path Wasserstein transfer bound with intrinsic squared error $m/n^2$, off-diagonal integrand estimates, and a weak-error transfer principle with additive error $n^{-1}$. Explicit models with feedback through both backward laws verify the cancellation assumptions. Examples distinguish the intrinsic backward error from the forward law error and establish matching lower bounds for each backward component.
Shu-Xian Gao, Ying Hu, Jia-Qiang Wen· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.