A family of bounded, globally Lipschitz denoisers for which both the forward-marginal error and the path-space total variation distance tend to zero, while their Euler--Maruyama endpoints diverge in every $W_p$ for compactly supported data.
Abstract
Score matching controls average error under the forward marginals, but a discretized reverse-time sampler evaluates the learned score along its own trajectory. We show that small forward-marginal error does not guarantee numerical stability. We construct a single smooth score field with arbitrarily small forward-marginal $L^2$ error. The learned reverse-time process is nonexplosive, has moments of every order, and can be arbitrarily close to the exact reverse-time process in path-space total variation. Yet its Euler--Maruyama discretizations converge in probability while every positive moment diverges. Thus weak convergence can hold even though every Wasserstein distance $W_p$, $p\ge1$, diverges. The same failure can occur within one fixed finite neural architecture. We construct a family of bounded, globally Lipschitz denoisers for which both the forward-marginal error and the path-space total variation distance tend to zero, while their Euler--Maruyama endpoints diverge in every $W_p$. For compactly supported data, we also give a simple positive result. Projecting the learned denoiser onto a known bounded closed convex set containing the support preserves pointwise accuracy, gives grid-uniform moment bounds, and yields Wasserstein convergence under mild local regularity. Experiments with a small fixed DiT-style network show large growth along rare numerical trajectories and its suppression by denoiser projection, while overall trajectory errors remain small.
This analysis transfers reverse-time discretization errors to the forward corruption law and treats two numerical schemes within a common framework, which identifies a local error, accumulates it through the forward evolution, and inserts the result into a common KL decomposition.
A 170M-parameter M2S model trained on about 262B OpenWebText token slots outperforms the evaluated pure-uniform SEDD, GIDD, and Neural CTMC checkpoints at every tested sampling budget, reaching generative PPL $143.3$ at 128 steps versus $183.6$ for the strongest pure-uniform baseline.
This work considers a first-order sampler based on the leave-one-out denoiser for uniform and remasking processes whose coordinate updates can be performed in parallel, and derives an exact information-theoretic representation of the discretization error in terms of the mutual information between different coordinates...
GeometricSPRINT (Geometric Step Pruning for Inference in Trajectories), a training-free framework for constructing non-uniform sampling schedules from the geometry of denoising trajectories, consistently improves over uniform DDIM (Denoising Diffusion Implicit Models) schedules at matched NFE budgets.
Direct first-order jackknife cancellation and exact-LOO concentration control deletion-to-full risk transfer and fluctuation, respectively, completing recovery of the conditional population-risk curve are estimated.
It is shown that the alternating COMB control attains the limiting Hamiltonian maximum exactly when the two highest scores coincide and the third- and fourth-highest scores coincide, and that the alternating COMB control attains the limiting Hamiltonian maximum exactly when the two highest scores coincide.
Erhan Bayraktar, Ibrahim Ekren, N. Kolliopoulos· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.