Skip to content

Semi-Supervised Conditional Generative Learning through Stochastic Interpolation and Sufficient Representations

Jul 2026 · arXiv.org · Vol abs/2607.16725 · 0 citations · 40 references
Mathematics Computer Science

TL;DR

The theoretical results demonstrate that this method effectively mitigates the curse of dimensionality inherent in direct ambient-space generative modeling and derive non-asymptotic convergence rates proving that RepG significantly improves sample complexity.

Abstract

Conditional generative modeling remains a challenging problem in semi-supervised settings where labeled data is scarce but unlabeled samples are abundant. To effectively leverage structural information embedded within the unlabeled dataset and compensate for sparse conditioning signals, we propose a semi-supervised framework combining conditional stochastic interpolation with low-dimensional latent representations. RepG decomposes generation into two stages: label-dependent latent sampling and high-dimensional reconstruction. This isolates the supervised learning of conditional dependencies to a low-dimensional space, requiring few labels while utilizing the abundant unlabeled data purely for reconstruction. Theoretically, we establish an error decomposition showing that the Kullback-Leibler divergence of RepG comprises stage-wise estimation errors and a structural bias quantified by conditional mutual information. For deep neural network estimators, we derive non-asymptotic convergence rates proving that RepG significantly improves sample complexity. By confining the supervised estimation burden to the low intrinsic dimension of the latent representation, RepG achieves a strictly faster convergence rate. Complemented by a minimax lower bound, our theoretical results demonstrate that this method effectively mitigates the curse of dimensionality inherent in direct ambient-space generative modeling.

View source

Similar papers

Jul 2026

Semi-Supervised Conditional Diffusion via Label Augmentation

This work introduces label-augmented conditional diffusion (LACD), a simple and effective approach that incorporates unlabeled examples by assigning them a designated trivial label and performing joint denoising score matching over the augmented dataset.

Jin Su, Yuan Gao, Yong Zhou et al. · 0 citations
#machine learning Preprint Aug 2026

Learning the Geometry of Admissible Hypotheses through Inductive Bias in Training Distributions

This work presents a framework for learning continuous latent representations of admissible partial differential equations by embedding a scientific inductive bias directly into the training distribution, and shows that embedding a scientific inductive bias in the training distribution enables the learning of compact a...

J. Crowley, Faez Ahmed, A. van Beek · 0 citations
Preprint Aug 2026

Conditional-Independence-Regularized Distributional Autoencoders for Mixed-Type Data

Mixed-type data containing both numerical and categorical variables arise in many scientific and real-world applications. Existing representation learning and generative modeling approaches typically focus either on reconstruction accuracy or unconditional data generation, but often fail to recover the full conditional...

Si-Yuan Tang, Gong-Jun Xu, Ji Zhu · 0 citations
Preprint Aug 2026

Geometric Regularization for Long-Tailed Semi-Supervised Learning via Gaussian Feature Bridges

This work introduces a novel framework, Gaussian Bridge Consistency (GBC), to address challenges of semi-supervised learning by constructing semantic interpolation paths between unlabeled samples and high-quality class anchors, and proposes BridgeMix, a confidence-aware feature mixing strategy that interpolates both sa...

Hong-Yang He, Xin-Yuan Song, Yan Zhong et al. · 1 citation
Open access Aug 2026

Weakly-Supervised Learning with Partial Multi-Labels by Leveraging Dual Label Correlation Perspectives

A novel PML method, namely Wasserstein Partial Multi-Label Learning with dual Label Correlation Perspectives (Wpml3cp), solved by the gradient descent with an augmented Lagrange multiplier technique, and empirical results demonstrate that Wpml3cp and Wpml3cp-D can outperform the PML baselines in various noisy levels.

Xi-Ming Li, Yuanchao Dai, Bing Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.