This work proposes a real-time video compression framework that represents and compresses a Group of Pictures using a coarse-to-fine multi-scale 2D Gaussian Splatting structure coupled with a lightweight deformation network and demonstrates the potential of Gaussian Splatting as a practical solution for efficient high-resolution video compression.
Abstract
Ultra-High-Definition (UHD) video presents significant challenges for efficient storage and real-time decoding. Learning-based methods, such as Neural Video Compression (NVC) and Implicit Neural Representations (INR), achieve competitive rate-distortion performance but suffer from high decoding latency and excessive memory usage. Meanwhile, Gaussian Splatting has recently attracted attention in the computer graphics community due to its ultra-fast rendering and high-fidelity visual quality. Despite these advantages, its application in video compression remains largely unexplored. To bridge this gap, we propose a real-time video compression framework that represents and compresses a Group of Pictures (GOP) using a coarse-to-fine multi-scale 2D Gaussian Splatting (2DGS) structure coupled with a lightweight deformation network. Experiments demonstrate that our method delivers rate-distortion performance in LPIPS that surpasses H.265 and other state-of-the-art learning-based video compression methods. Our work demonstrates the potential of Gaussian Splatting as a practical solution for efficient high-resolution video compression.
Quantized Anchored Residual Coding Gaussian Streaming (QuARC-GS), a quantization-aware 4D scene optimization framework for online dynamic scene reconstruction that achieves ultra-high compression while maintaining reconstruction speed and quality, is proposed.
V. Nguyen, Yu-Chen Wang, Kyung Chul Lee et al.· 0 citations
A Video Compression framework built upon a foundational flow model that enables the compressor to harness generative video flow priors effectively, which reduces bit consumption by 58\% and achieves one-step decoding and reconstructions with high perceptual fidelity.
Yichong Xia, Qin-Hong Wu, Bin Chen et al.· 1 citation· ⚡1
This paper presents UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing and introduces Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequ...
3D Gaussian Splatting (3DGS) enables high-quality novel-view synthesis but requires substantial storage. Existing compression methods often rely on spatial context modeling over irregular 3D representations, increasing the complexity of training and coding. Meanwhile, floating-point context inference can introduce nume...
Peng-Peng Yu, Yue-Ru Chen, Fei Song et al.· 0 citations
This paper proposes BinRVR, a binarized RAW video restoration framework that reduces computation and parameters by approximately 96% while incurring only about 4% performance degradation, and develops a Distribution-Aware Binarized Convolution (DAB-Conv) that leverages the statistics of full-precision activations to mi...
Tianyu Zhu, Ying Fu, He-Song Li et al.· IEEE Transactions on Pattern...· 0 citations
Immersive video communication requires photorealistic, render-efficient, and compact dynamic scene representations. 3D Gaussian Splatting (3DGS) offers a promising representation, but dynamic 3DGS remains difficult to compress due to dense primitives and spatiotemporal redundancy. Anchor-based formulations improve comp...
Ge Gao, Si-Yue Teng, Chanqgi Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.