Skip to content

FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers

Jul 2026 · arXiv.org · Vol abs/2607.14898 · 1 citation · ⚡ 1 influential · 81 references
Computer Science

TL;DR

FlashDecoder is introduced, a fast, memory-efficient pure-Transformer video decoder that decodes latents to pixels frame by frame, enabling constant-latency streaming in real-time video generation.

Abstract

Real-time video generation demands fast decoding as much as fast denoising, yet current latent video diffusion models rely on 3D convolutional decoders that are slow and memory-intensive at high resolutions or for long video. We introduce FlashDecoder, a fast, memory-efficient pure-Transformer video decoder that decodes latents to pixels frame by frame. At each step, the current frame attends only to a fixed-size window of past frames through a rolling KV cache. The fixed temporal window keeps decoding fast and memory bounded regardless of video length, enabling constant-latency streaming. Because frames are processed sequentially, temporal causality is enforced without explicit attention masks, enabling training at resolutions up to 1080p and matching the reconstruction quality of convolutional decoders. On the Wan2.1 and Wan2.2 latent spaces, FlashDecoder matches each convolutional decoder in reconstruction quality (e.g., 41.55dB vs. 41.49dB PSNR at 1080p) while decoding 3.6x-4.7x faster with up to 11x less memory on a single H100 GPU. With architecture-aware inference optimizations, the speedup widens to 12x.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

VoRTeC: Taming Foundation Flow for One-step Real time Video Compression

A Video Compression framework built upon a foundational flow model that enables the compressor to harness generative video flow priors effectively, which reduces bit consumption by 58\% and achieves one-step decoding and reconstructions with high perceptual fidelity.

Yichong Xia, Qin-Hong Wu, Bin Chen et al. · 1 citation · ⚡1
Preprint Sep 2026

FastVR: Efficient Streaming Video Restoration with One-Step Diffusion

Diffusion-based video restoration recovers realistic details, but its practical deployment is limited by two efficiency bottlenecks: costly VAE encoding and decoding, and the quadratic cost of full self-attention in diffusion transformers (DiTs). This paper presents FastVR, a streaming video restoration framework built...

Xiao-Xu Chen, Qin Yang, Hao-Ran Bai et al. · 0 citations
Preprint Sep 2026

tcnerv:dual-domain temporal context modeling for implicit neural video compression

Video compression aims to minimize reconstruction distor tion under a constrained bit rate. Existing video implicit neural representations (INRs) often decode frames independently, leaving intermediate features unconditioned on previous reconstructions and content embeddings without explicit temporal prediction. We pro...

Xue-Lian Xiang, Yi-Xin Zhao, He-Qi Xiang et al. · 0 citations
Preprint Aug 2026

DiffVC-ONE: Diffusion-based Generative Video Compression with One-Step Video Diffusion Transformer

DiffVC-ONE, a diffusion-based generative video compression framework built on a one-step Video Diffusion Transformer, is proposed and a Unified Unidirectional Latent Compressor that uses a shared model to efficiently and uniformly compress compact latent slices is introduced.

Wenzhuo Ma, Zhen-Zhong Chen · 0 citations
Preprint Sep 2026

Streaming4D: Accelerate 4D World Models via Block-wise Video Generation and Incremental Reconstruction

Current 4D generation paradigms are often bottlenecked by a sequential decoupling design: video is generated first, followed by 3D reconstruction, leading to high interaction latency. This limits applications in interactive real-time scenarios. To this end, we propose \textbf{Streaming4D}, a tightly coupled synchronous...

Xiao-Yan Liu, Jia-Xin Liu, Kang Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for Video World Models

Video world models achieve long-range temporal consistency by storing KV cache during generation, but the growing cache makes KV cache memory a major deployment bottleneck, which motivates low-bit quantization study for efficiency. Existing 2-bit KV cache quantization methods can achieve nearly lossless performance on...

Jia-Qi Zhao, Xiao-Bin Hu, Bo Yin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.