Few-step autoregressive video generation enables efficient streaming synthesis, but errors introduced in early temporal blocks are reused as context and can propagate through subsequent rollouts, leading to detail degradation, structural drift, and unstable motion. Existing distribution matching distillation (DMD) prim...
Fang-Yu Lin, Xing-Tong Ge, Lu Zhu et al.· 0 citations
Omni-LiveAvatar is presented, the first framework for minute-level, real-time streaming joint audio-video avatar generation and proposes a progressive autoregressive distillation pipeline that transfers a large bidirectional joint audio-video diffusion model into a few-step autoregressive generator without auxiliary st...
Lu Zhu, Xing-Tong Ge, Fang-Yu Lin et al.· 1 citation
Mage-VL is presented, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction and establishes AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes.
The results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.