Skip to content

BLUE: Semantics-Preserving Video Compression for Efficient Vision-Language Surveillance Analytics

Jul 2026 · arXiv.org · Vol abs/2607.19515 · 0 citations · 22 references
Engineering Computer Science

TL;DR

It is suggested that BLUE functions as a machine-centric compression layer for surveillance video, reducing bandwidth and inference cost while preserving VLM semantic performance.

Abstract

Continuous surveillance video creates a growing storage, transmission, and inference burden for enterprise video analytics systems. While modern codecs such as H.265 reduce bitrate for human-viewable video, aggressive compression can degrade downstream computer-vision performance and does not necessarily reduce the number of vision-language model (VLM) inference calls required for semantic video understanding. This paper evaluates BLUE, a fixed-camera surveillance compression approach that suppresses static-background redundancy while preserving foreground activity, for its effect on VLM-based event and anomaly understanding. We compare raw H.265 and BLUE-compressed H.265 video on two surveillance datasets: VIRAT, comprising 227 paired event samples from 106 clips, and CHAD, comprising 54 human-activity anomaly clips. For each pair, the same frame index is evaluated using a VLM captioning pipeline, and outputs are scored against annotation-derived ground truth using a blind judging protocol. The results show no measurable degradation in semantic inference quality. On VIRAT, the mean VLM score remains effectively unchanged between raw H.265 and BLUE, with a mean difference of approximately -0.01 on a 0-10 scale. On CHAD, raw H.265 and BLUE obtain near-equivalent mean scores of 4.31 and 4.26, respectively. Compression saving is also uncorrelated with VLM score change on VIRAT (r = 0.004), indicating that higher BLUE compression does not predict semantic quality loss. Beyond storage reduction, BLUE increases the share of skip-heavy P-frames on CHAD from 1.4% to 53.2%, enabling an estimated 53% reduction in VLM calls through packet-size-based frame skipping. These findings suggest that BLUE functions as a machine-centric compression layer for surveillance video, reducing bandwidth and inference cost while preserving VLM semantic performance.

View source

Similar papers

#computer vision Review Sep 2026

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal grounding comes at a co...

Killian Steunou, Yannis Tevissen, M. E. El Yacoubi · 0 citations
Conference Aug 2026

Distilled Vision-Language Model Semantics for Edge-Deployable Content-Aware Adaptive Video Compression

With the proliferation of edge video services, content-aware adaptive video compression has become increasingly critical to balancing bandwidth efficiency and visual quality. However, conventional codec control methods mainly rely on low-level signal statistics, limiting their adaptability to high-level content semanti...

Qian Wei, En-Fang Cui, Zhi-Yuan Liang et al. · 0 citations
Preprint Sep 2026

SemABR: Measuring Video Semantic Fidelity with Multimodal LLMs for Adaptive Bitrate Streaming

Conventional video metrics such as PSNR, SSIM, and VMAF measure visual distortion or perceptual quality, but they do not directly capture semantic preservation: whether compression retains a video's objects, actions, and temporal narrative. Existing Quality-of-Experience (QoE)-driven bitrate-selection and resource-allo...

Shi-Qi Xu, S. Liew, Yu-Yang Du · 0 citations
#computer vision Preprint Sep 2026

ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding

ShallowStream is a novel framework that leverages the shallow layers of an MLLM to simultaneously perform frame encoding and retrieval index building and achieves performance on par with the strongest existing streaming methods, while reducing per-frame prefill latency and 10-second end-to-end latency by up to 52.1x an...

Jitai Hao, Ke-Shuai Yang, Qiang Huang et al. · 0 citations
Preprint Aug 2026

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models

This work proposes a two-stage adaptive token pruning strategy specifically designed for video processing that improves accuracy by +7\% on a video captioning benchmark at 10% token retention, while reducing computation TFLOPs by 95\%.

Paribesh Regmi, Qingshuang Chen, Chi Zhang et al. · 1 citation
Preprint Aug 2026

Think in Sets for Streaming Video Token Compression

Streaming VideoLLMs process frames causally while visual tokens grow continuously, making compression essential for controlling prefilling latency and memory. Existing training-free methods independently rank tokens, ignoring marginal-gain interactions among retained tokens. We argue that streaming video token compress...

Moxu Duan, Jingwen Fu, Yuwang Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.