Skip to content

ReGenVC: End-to-End Real-Time Generative Video Coding at Ultra-Low Bitrate

Jul 2026 · arXiv.org · Vol abs/2607.28144 · 0 citations · 42 references
Computer Science Engineering

TL;DR

ReGenVC is the first end-to-end generative video codec to combine ultra-low-bitrate encoding with real-time decoding on an 8-GPU system and three model-preserving system techniques are presented.

Abstract

We present ReGenVC, an end-to-end generative video codec that compresses talking-head video to an ultra-low bitrate and decodes it in real time. The encoder reduces a source clip to a compact bitstream -- a neurally compressed first frame, per-frame pose keypoints, and metadata -- totaling about 26 kB for a 77-frame sequence. The decoder is a four-step distilled diffusion transformer that reconstructs the video conditioned on the transmitted pose and reference frame. Compared with x264/x265, ReGenVC reduces the bitrate to roughly one tenth of that required by traditional codecs (about 26 kB vs. 250--280 kB for essentially artifact-free reconstruction); at a matched ultra-low bitrate, conventional codecs collapse into blocking artifacts while ReGenVC stays sharp by exploiting a strong generative prior. The central obstacle to deploying such a codec is decoder latency: multi-step sampling with transformer and VAE components is too slow for interactive use. We make the decoder real-time through four-step distillation and three model-preserving system techniques: (i) eight-GPU unified sequence parallelism (Ulysses&Ring), (ii) a spatially-split VAE, and (iii) a three-stage overlapped pipeline; an analytical timing model characterizes the real-time feasibility region. On an 8-GPU node, the system sustains 24 fps output (972 ms per 25-frame window, within the 1000 ms budget), enabling a live browser stream without observed frame underruns. A hybrid CPU-GPU deployment further runs the encoder on the CPU at 24 fps and offloads the decoder-side one-shot conditioning encoders to the CPU, reducing the per-GPU memory peak from 21.1 GB to about 7.7 GB. To our knowledge, ReGenVC is the first end-to-end generative video codec to combine ultra-low-bitrate encoding with real-time decoding on an 8-GPU system.

View source

Similar papers

Preprint Aug 2026

GVC-RT: Towards Real-Time Generative Video Compression at Ultra-Low Bitrates

This work systematically identifies the computational bottlenecks and proposes GVC-RT, which redesigns the generative latent coding framework to realize real-time video coding without sacrificing compression performance, and introduces a lightweight de-tokenizer architecture to resolve the final latency bottleneck duri...

Tian-Jian Dang, Si-Xian Wang, Lei Luo et al. · 1 citation · ⚡1
Preprint Aug 2026

GVCCTurbo: Rate-Compute Quality Scheduling for Codebook Driven Generative Compression

This work proposes GVCCTurbo, a BPP-driven scheduler that separates expensive prior refreshes from codebook corrections, and supports BPP-to-compute scheduling as a controllable extension of sampler-length tuning, without requiring the allocated point to dominate every boundary point.

Ziyue Zeng, Ding-Jie Peng, Xun Su et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding

Neural audio codecs are a fundamental component of modern speech generation systems. While recent codecs achieve increasingly low bitrates, reducing frame rate remains challenging, as each token must preserve more information while maintaining reconstruction quality. We present ZipCodec, a streaming neural speech codec...

Luca Della Libera, Cem Subakan, M. Ravanelli · 0 citations
Open access Jul 2026

LVT: A Learned Video Transcoding Framework

A learned video transcoding framework (LVT) is proposed to optimize video transcoding, leveraging coding priors from the input bitstream to guide the transcoding process, and outperforms both existing traditional and learned video codecs in transcoding performance.

Nian-Xiang Fu, Dai-Qin Yang, Zhe-Nan Lin et al. · 0 citations
Jul 2026

FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers

FlashDecoder is introduced, a fast, memory-efficient pure-Transformer video decoder that decodes latents to pixels frame by frame, enabling constant-latency streaming in real-time video generation.

Minguk Kang, Suha Kwak · 1 citation · ⚡1
Preprint Sep 2026

Scalable Neural Video Representation Compression

Scalable video coding (SVC) encodes a video into a layered bitstream consisting of a base layer and one or multiple enhancement layers, enabling decoding at different bitrate/quality/resolution operating points to accommodate diverse device capabilities and network conditions. Due to its practical flexibility, SVC has...

Tian-Hao Peng, H. Kwan, Fan Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.