Skip to content

Semi-Supervised Conditional Diffusion via Label Augmentation

Jul 2026 · arXiv.org · Vol abs/2607.16685 · 0 citations · 63 references
Computer Science Mathematics

TL;DR

This work introduces label-augmented conditional diffusion (LACD), a simple and effective approach that incorporates unlabeled examples by assigning them a designated trivial label and performing joint denoising score matching over the augmented dataset.

Abstract

Conditional diffusion models have become a powerful and flexible framework for learning complex conditional distributions from labeled data. In practice, however, acquiring high-quality labels is costly and time-consuming, leaving large volumes of unlabeled data unused. To address this, we introduce label-augmented conditional diffusion (LACD), a simple and effective approach that incorporates unlabeled examples by assigning them a designated trivial label and performing joint denoising score matching over the augmented dataset. We provide sufficient conditions guaranteeing population-level identifiability of the target conditional distribution under this scheme. Moreover, we establish rigorous statistical guarantees: when sufficiently many unlabeled samples are available, the sampling distribution produced by LACD converges strictly faster than the purely supervised estimator in total variation distance, and at least as fast in Wasserstein-1 distance. Extensive experiments on synthetic, image, and tabular benchmarks corroborate our theory and show substantial gains in sample efficiency and generative performance compared with the purely supervised estimator.

View source

Similar papers

Preprint Aug 2026

Geometric Regularization for Long-Tailed Semi-Supervised Learning via Gaussian Feature Bridges

This work introduces a novel framework, Gaussian Bridge Consistency (GBC), to address challenges of semi-supervised learning by constructing semantic interpolation paths between unlabeled samples and high-quality class anchors, and proposes BridgeMix, a confidence-aware feature mixing strategy that interpolates both sa...

Hong-Yang He, Xin-Yuan Song, Yan Zhong et al. · 1 citation
Open access Aug 2026

Weakly-Supervised Learning with Partial Multi-Labels by Leveraging Dual Label Correlation Perspectives

A novel PML method, namely Wasserstein Partial Multi-Label Learning with dual Label Correlation Perspectives (Wpml3cp), solved by the gradient descent with an augmented Lagrange multiplier technique, and empirical results demonstrate that Wpml3cp and Wpml3cp-D can outperform the PML baselines in various noisy levels.

Xi-Ming Li, Yuanchao Dai, Bing Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Semi-Supervised Learning under Spatially Biased Sampling

It is demonstrated that spatial non-stationarity contributes to performance loss independently of marginal mismatch and that models become increasingly overconfident outside the regions where labels are available, and introduced a kernel-weighted local divergence metric that provides a more stable estimate of spatial m...

Bright Wiredu Nuakoh, F. Fouedjio, Stephen Bradshaw et al. · 0 citations
#machine learning Preprint Sep 2026

Improved Distributional Diffusion Models

Distributional Diffusion Models (DDMs) replace the standard mean-prediction denoiser with a \emph{distributional} denoiser trained via a scoring rule objective, learning a stochastic approximation to $p(x_1 \mid x_t)$ rather than its conditional mean. However, scaling DDMs to modern image-generation settings faces two...

Tommaso Martorella, Alexandre Galashov, F. Krause et al. · 1 citation
#machine learning Preprint Sep 2026

Robustness of Diffusion Models under Distribution Shift

Score-based diffusion models are increasingly considered in settings where the underlying data distribution may differ from the training distribution, yet existing theoretical guarantees largely focus on the no-shift setting. In this work, we study robust score estimation under Wasserstein perturbations of a reference...

Wei Luo, N. K. Chada, Shi-Jie Zhang et al. · 0 citations
Open access 2026

Diffusion-Based Generative Regularization for Cross-Modal Supervised Learning

Diffusion-based generative regularization is proposed, a supervised discriminative learning framework that leverages a frozen diffusion-based generative model as a regularizer without explicitly generating additional training samples to improve supervised discriminative learning.

Takuya Asakura, Nakamasa Inoue, Koichi Shinoda · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.