\emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce \emph{a unified theoretical framework} that separates distribution modeling from matching discrepancy and connects global objectives to poin...
Chi Zhang, Hao-Yan Shi, Yue-Yi Liu et al.· 1 citation
Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework mi...
Chi Zhang, Yue-Yi Liu, Hao-Yan Shi et al.· 1 citation
Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatars, and embodied agents. While recent audio-video generative models can synthesizehigh-fidelity synchronized content, existing supervision is...
Chi Zhang, Hao-Yan Shi, Yueyi Liu et al.· 2 citations
This work introduces MeetingToM, a benchmark for complex social behavior reasoning in naturalistic multi-party meetings and establishes MeetingToM as a testbed for advancing meeting-grounded ToM in multimodal models.