Skip to content
Open access

Why diffusion models do not memorize: the role of implicit dynamical regularization in training

Aug 2026 · Journal of Statistical Mechanics: Theory and Experiment · Vol 2026 · 2 citations · 66 references
Physics

TL;DR

It is found that τmem increases linearly with the training set size n, whereas τgen remains constant, which creates a growing window of training times with n where models generalize effectively, despite showing strong memorization if training continues beyond it.

Abstract

Diffusion models have attained remarkable success across a wide range of generative tasks. A key challenge lies in understanding the mechanisms that prevent their memorization of training data and allow generalization. In this work, we investigate the role of the training dynamics in the transition from generalization to memorization. Through extensive experiments and theoretical analysis, we identify two distinct timescales: an early time τgen at which models begin to generate high-quality samples and a later time τmem beyond which memorization emerges. Crucially, we found that τmem increases linearly with the training set size n, whereas τgen remains constant. This creates a growing window of training times with n where models generalize effectively, despite showing strong memorization if training continues beyond it. It is only when n becomes larger than a model-dependent threshold that overfitting disappears at infinite training times. These findings reveal a form of implicit dynamical regularization in the training dynamics, which allows to avoid memorization even in highly overparameterized settings. Our findings are supported by numerical experiments with standard U-Net architectures on realistic and synthetic datasets, alongside a theoretical analysis using a tractable random features model studied in the high-dimensional limit.44 https://github.com/tbonnair/Why-Diffusion-Models-Don-t-Memorize. https://github.com/tbonnair/Why-Diffusion-Models-Don-t-Memorize.

Read PDF

Similar papers

Preprint Aug 2026

Tunneling the Loss Landscape: Bypassing Memorization with Monte Carlo Parameter Swapping

Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization. While previous works have attempted to interpret it through classical machine learning mechanisms like weight norm, recent research draws an analogy from statistical physics, framing grokking as a form of computational glass relaxation. This theory defines the initial memorization as a result of `fast cooling'where the training loss is reduced so quickly that a glass state is formed, followed by a `slow relaxation'towards final generalization. Although providing a unifying framework for representative grokking theories, this perspective has remained largely at the theoretical on macroscopic level without direct empirical validation on training dynamics. Here we introduce a three-component framework to directly characterize the training dynamics via parameter mobility (PM), and two representative measurements from glassy dynamics: replica correlation (RC) and fractal dimension (FD). We demonstrate that standard optimization presents clear signatures of glass dynamics and inherently traps the grokking network in a kinetic arrested memorization state with a collapsed mobility, strong history dependence, and channel-like motions. This quantitative agreement motivates us to introduce State-Aware Monte Carlo Parameter Swapping (SAM-Swap), an optimization plug-in that can accelerate generalization, inspired by swap Monte Carlo algorithm widely used in glass dynamics. Comparing SAM-Swap, weight decay, and Gaussian gradient noise, we find that accelerated generalization is consistently associated with random exploration in the parameter space, similar to diffusion in physics.

Lai Shun Chan, Xiaotian Zhang, Yue Shang et al. · 0 citations
Preprint Aug 2026

Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime

This work develops a generative counterpart to the theory of benign overfitting and algorithmic regularization for overparameterized neural networks in the supervised lazy-training regime by studying denoising score matching in a vector-valued reproducing kernel Hilbert space with an inner-product kernel.

Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian et al. · 0 citations
Jul 2026

The Grokked Illusion: True Equilibrium Mitigates Catastrophic Forgetting

While neural networks are typically evaluated by their training and test performance, these metrics do not reveal how robust a learned representation is. Recent studies have shown that solutions occupying larger volumes in parameter space, as quantified by Boltzmann entropy, often exhibit superior generalizability compared to those reached by conventional optimization, a phenomenon known as the high entropy advantage. Here we ask whether this advantage persists beyond generalization. Specifically, we investigate models'robustness, the ability to retain the learned knowledge when the model is subsequently trained to acquire new information. Using grokking in modular arithmetic as a controlled setting, we design a noise injection experiment to evaluate the robustness difference between AdamW-trained transformers and high-entropy model sampled from Wang-Landau Molecular Dynamics with identical saturated performance. By forcing both models to fully remember new data with random labels, we find that AdamW-trained models suffer from catastrophic forgetting, with original task test accuracy dropping from 100% to below 75%, whereas the high-entropy models maintain approximately 95% test accuracy. We term this hidden fragility behind apparent generalization the"grokked illusion."Through singular value decomposition of the neural network weights, we discover that high-entropy neural networks possess significantly higher effective rank in attention and MLP layers both before and after noise injection, indicating richer feature representations can serve as a buffer against catastrophic forgetting. Our findings demonstrate that perfect generalization does not imply equal robustness, offering a new perspective on what makes a trained model robust to interference.

Xiaotian Zhang, Lai Shun Chan, Yue Shang et al. · 0 citations
#machine learning Preprint Aug 2026

Canalization Before Generalization: Grokking as a Dynamical Probe

For overparameterized neural networks, many solutions can fit the training data equally well while behaving very differently on unseen samples. Grokking separates training fit from visible generalization, providing a window for studying how this selection develops during training. We sweep short, fixed-duration weight-decay (WD) perturbations across the pre-generalization plateau and measure how they shift later generalization time. Across three grokking tasks, these shifts are unordered early in the plateau but later form a stable dose ordering, with stronger WD increases leading to earlier generalization and stronger WD decreases leading to later generalization. This ordering emerges before visible generalization in all three tasks. Meanwhile, test-loss barriers between perturbed and baseline generalization checkpoints collapse toward zero while the ordered timing effects persist. Drawing on Waddington's developmental landscape as an analogy, we call this combination of increasingly constrained solution selection and persistent dose-ordered shifts in generalization timing the canalization of grokking solution selection.

Yi-Min Lin · 0 citations
Preprint Aug 2026

Continual-learning rules shape representational drift

Together, these results link representational drift to the stability--plasticity trade-off: its magnitude is shaped by the mechanism that protects old knowledge, and suppressing it can restrict future learning.

Yi-Kai Si, Shanshan Qin · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.