Distilled Vision-Language Model Semantics for Edge-Deployable Content-Aware Adaptive Video Compression
Abstract
With the proliferation of edge video services, content-aware adaptive video compression has become increasingly critical to balancing bandwidth efficiency and visual quality. However, conventional codec control methods mainly rely on low-level signal statistics, limiting their adaptability to high-level content semantics. Although vision-language models (VLMs) offer rich semantic understanding, their prohibitive inference costs preclude direct deployment in resource-constrained edge environments. To bridge this gap, this paper proposes CAVC, an edge-deployable content-aware adaptive video compression framework leveraging distilled VLM semantics. Specifically, semantic knowledge from a large VLM is distilled into a lightweight video feature extractor, yielding compact content representations that inform adaptive compression decisions. A soft actor-critic policy is then trained to optimize compression parameters under quality constraints. Experiments on UVG and MCL-JCV show that the proposed method improves rate-distortion efficiency over x265 and achieves competitive performance compared with representative learned compression baselines, while requiring substantially lower deployment overhead than end-to-end learned codecs. These results demonstrate the practicality of CAVC for resource-constrained edge video scenarios.