Non-linear Data-driven Transforms and Intra Prediction Modes for Block-based Video Compression.
Abstract
Combining the benefits of predictive coding and energy-compacting transforms is a common method in visual data compression. For instance, the video coding standard Versatile Video Coding (VVC) supports different intra-prediction modes for predicting a block of original samples. Furthermore, the codec has various transform options: trigonometric ones like the DCT-II and non-separable ones named LFNST. The transform coefficients are quantized and entropy coded, and the encoder can evaluate the rate-distortion impact of the quantization decisions in several ways. However, caused by the emergence of efficient learned image codecs, there is growing interest in training non-linear transforms for block-based video coding. It has already been shown that non-linear modifications of existing block transforms in VVC yield additional coding gain. This paper extends the approach to jointly optimizing transforms and intra modes with a new training method. Here, the learned transforms use an orthogonal transform matrix combined with a neural network for computing a non-linear correction term, while the learned intra modes have a similar network design. At the training stage, Trellis-Coded Quantization (TCQ) is partly applied to the transform coefficients: only the quantization to the zero level is performed and larger coefficients remain unchanged. Furthermore, a non-linear bitrate model is fitted to the empirical distribution of the transformed residual for estimating the bitrate impact of the learned analysis transform. Over the VTM-14.2 anchor, the average bitrate savings of the learned intra modes and transforms range from 0.8 % to 0.9 % across different classes. Importantly, the learned transforms do not depend on the reference samples as input.