Skip to content
Preprint

Omni-LiveAvatar: Minute-Level Real-Time Streaming Joint Audio-Video Avatar Generation

Aug 2026 · 0 citations · 36 references
Computer Science

TL;DR

Omni-LiveAvatar is presented, the first framework for minute-level, real-time streaming joint audio-video avatar generation and proposes a progressive autoregressive distillation pipeline that transfers a large bidirectional joint audio-video diffusion model into a few-step autoregressive generator without auxiliary stabilization mechanisms.

Abstract

Joint audio-video generative models serve as foundation for immersive and interactive digital-human generation. Nevertheless, most existing models rely on bidirectional attention and multi-step denoising and can generate only short clips, making them unsuitable for real-time interaction over extended durations. We present Omni-LiveAvatar, the first framework for minute-level, real-time streaming joint audio-video avatar generation. Specifically, we propose (1) a progressive autoregressive distillation pipeline that transfers a large bidirectional joint audio-video diffusion model into a few-step autoregressive generator without auxiliary stabilization mechanisms; (2) a synchronized audio-video long-short-term memory that preserves global consistency under a bounded memory budget; and (3) a hierarchical rolling prompt planning strategy that enables coherent semantic evolution and seamless prompt transitions. Extensive experiments show that Omni-LiveAvatar generates high-quality, synchronized minute-level avatars in real time. In terms of speed, it achieves a 33$\times$ generation speedup over its teacher, LTX-2, on a single NVIDIA H200 GPU; in terms of generation quality, it outperforms accelerated baselines across visual quality, audio quality, cross-modal synchronization, and human fidelity. Our code is available at https://github.com/Aoko955/Omni-LiveAvatar.

View source

Similar papers

Jul 2026

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars

OmniMate, a unified framework for open-ended real-time interactive audio-visual avatar generation and a Multi-Reference Conditioning Module (MRCM), which leverages multiple reference images and a reference speech segment to provide persistent visual and speaker identity cues throughout long-duration streaming interactions.

Quan-Yue Song, Yi-Shan He, Yan-Bo Ding et al. · 0 citations
Preprint Aug 2026

LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time

This work presents LiveAnimate, to their knowledge the first animation system to combine real-time streaming with stable long-form generation at billion scale, built on a 14B-parameter video Diffusion Transformer (DiT).

Yuxuan Zhang, H. Xiong, Yubo Huang et al. · 0 citations
Jul 2026

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars

Production-ready audio-driven avatar generation requires efficient inference without sacrificing fidelity or motion expressiveness. However, existing acceleration methods often compromise quality through restrictive architectural choices, such as causal attention and short temporal horizons, or by reducing model capacity and resolution. Without such compromises, we propose AptAvatar, a 14B-parameter long-form audio-driven avatar generation framework that delivers fast and expressive inference. For efficiency in production-level applications, AptAvatar addresses the extreme two-step generation challenge. To bridge the gap between the multi-step teacher model and the two-step student model, we introduce Endpoint-Anchored Distribution Distillation. It augments vanilla distribution matching with a dedicated Anchor Score Estimator trained on the trajectory-endpoint distribution defined from a frozen pretrained 4-step bridge generator. This provides an attainable endpoint-level anchor for the evolving two-step student. To improve long-horizon consistency, we further introduce Self-Generated History Replay, which reuses cached outputs from earlier generator checkpoints as history conditions during chunk-wise training. This approximates inference-time conditioning on self-generated histories without costly online rollouts, mitigating quality degradation from accumulated history errors. Extensive experiments demonstrate that AptAvatar generates vivid 720p long-form avatar videos with only 2 NFEs, achieving a 60x speedup while preserving visual fidelity and long-horizon identity. Code is available at https://github.com/TaoLiveAIGC/AptAvatar

Hengyuan Zhang, Jing-Na Sun, Mei-Guang Jin et al. · 1 citation
Preprint Aug 2026

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

The method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation, and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift.

Yicheng Xiao, Wenxun Dai, Xinran Qin et al. · 2 citations
Jul 2026

Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory

Ripple is a real-time joint audio-video generation system with a cross-modal recurrent memory mechanism that combines a fixed-length sliding-window attention with modality-specific memory states that continuously summarize audio and video context.

Yan-Bo Ding, Zhi-Zhi Guo, Quan-Yue Song et al. · 0 citations
Preprint Aug 2026

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming

Vorch-Streamer, a post-training framework that addresses real-time long-form Text-to-Audio-Video (T2AV) streaming and enables real-time long-form avatar audio-video streaming, and bounded causal context and four-step denoising are presented.

Menglin Han, Yang Ding, Yulei Lu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.