Context-Aware Semantic Video Coding via Content-Adaptive Parameter Overfitting
Abstract
Hybrid semantic video coding combines learned representations with standardized residual coding, but existing approaches either retrain an entire model for each group of pictures or use a fixed decoder that cannot adapt to local content. This paper introduces a context-aware framework that specializes a pretrained semantic decoder to each group of pictures by optimizing only its channel-wise scale factors and transposed-convolution biases. The compact update is differentially compressed and transmitted with the semantic representation, enabling the receiver to reproduce the adapted decoder without full-model retraining. A hierarchical bidirectional prediction structure supplies motion-aligned temporal context, while a standardized residual pathway preserves reconstruction fidelity. Unlike prior hybrid semantic codecs, the proposed method jointly provides lightweight content adaptation, explicit accounting of parameter-transmission cost, and compatibility with a shared pretrained backbone. Experimental evaluation across natural and out-of-distribution video demonstrates consistent coding gains over the fixed-decoder baseline and the standardized reference codec under diverse conditions.