Skip to content
Open access

LVT: A Learned Video Transcoding Framework

Jul 2026 · ACM Trans. Multim. Comput. Commun. Appl. · Vol 22, pp. 1 - 23 · 0 citations · 40 references
Computer Science

TL;DR

A learned video transcoding framework (LVT) is proposed to optimize video transcoding, leveraging coding priors from the input bitstream to guide the transcoding process, and outperforms both existing traditional and learned video codecs in transcoding performance.

Abstract

With the exponential growth of video traffic and the continuous evolution of video coding standards, video transcoding has become essential for existing bitstreams to benefit from the advanced features of new video compression technologies. Typically, video transcoding involves decoding an existing bitstream and re-encoding the decoded sequence into a target format. A key challenge in transcoding is the inevitable presence of compression artifacts in the decoded sequences, which, if not properly addressed, can degrade transcoding efficiency by causing suboptimal bit allocation and disrupting core coding processes. In this article, a learned video transcoding framework (LVT) is proposed to optimize video transcoding, leveraging coding priors from the input bitstream to guide the transcoding process. In the framework, to mitigate the adverse effects of compression artifacts, a Coding Priors-Guided Spatial Feature Transform module is designed, which utilizes coding prior features to adaptively modulate intermediate features through spatial affine transformations, enhancing bit allocation and suppressing artifacts. Additionally, a Coding Priors-Guided Quality Adapter module is proposed to generate a compression degradation representation using coding priors, which dynamically interacts with intermediate features to enable the network to perceive and adapt to different levels of degradation in the input video. Furthermore, a Motion Vectors-Guided Flow Refinement module is proposed to reduce prediction errors caused by artifacts. It refines optical flow predictions by using motion vectors from the bitstream as auxiliary information. Extensive experiments demonstrate that our framework outperforms both existing traditional and learned video codecs in transcoding performance, achieving an average bitrate saving of 20.3% compared to the H.266/VVC reference software VTM under the practical YUV420 setting measured with PSNR.

Read PDF

Similar papers

Preprint Sep 2026

Scalable Neural Video Representation Compression

Scalable video coding (SVC) encodes a video into a layered bitstream consisting of a base layer and one or multiple enhancement layers, enabling decoding at different bitrate/quality/resolution operating points to accommodate diverse device capabilities and network conditions. Due to its practical flexibility, SVC has...

Tian-Hao Peng, H. Kwan, Fan Zhang et al. · 0 citations
Open access Aug 2026

MDFI: A Multi-Domain Features Integration for Compressed Video Quality Enhancement

This work proposes MDFI (Multi-Domain Features Integration), a compressed video quality enhancement approach that features a novel Frame-Prediction Feature Transform (FPFT) module to process prediction information to enhance decoded video quality.

Sang NguyenQuang, Hieu Bui Minh, BuiDinh Dang et al. · 0 citations
Preprint Aug 2026

Bit Allocation Transfer for Perceptual Quality Enhancement of Traditional Video Codecs

Traditional block-based video codecs, such as H.264/AVC, H.265/HEVC and H.266/VVC, rely on hand-crafted Rate-Distortion Optimization (RDO) processes that primarily minimize Mean Squared Error (MSE), which correlates poorly with human perceptual quality. While neural video compression methods can easily optimize percept...

Runyu Yang, Ivan V. Baji'c · 0 citations
Open access 2026

Context-Aware Semantic Video Coding via Content-Adaptive Parameter Overfitting

Hybrid semantic video coding combines learned representations with standardized residual coding, but existing approaches either retrain an entire model for each group of pictures or use a fixed decoder that cannot adapt to local content. This paper introduces a context-aware framework that specializes a pretrained sema...

P. Samarathunga, Yasith Ganearachchi, Thanuj Fernando et al. · 0 citations
Review Open access Aug 2026

A survey of implicit neural representations for video compression

Neural Video Representations (NVRs) have recently been proposed as a novel approach to the video compression problem. NVRs consist of one or multiple small neural networks that are overfitted on one specific video sequence, thereby encoding the video within the weights and biases of the network(s). In contrast to other...

Hannes Keunen, Maarten Wijnants, J. Liesenborgs · 2 citations
#artificial intelligence Preprint Sep 2026

VoRTeC: Taming Foundation Flow for One-step Real time Video Compression

A Video Compression framework built upon a foundational flow model that enables the compressor to harness generative video flow priors effectively, which reduces bit consumption by 58\% and achieves one-step decoding and reconstructions with high perceptual fidelity.

Yichong Xia, Qin-Hong Wu, Bin Chen et al. · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.