Skip to content

Optimizing Fidelity-Perception Tradeoff via Large Vision–Language Model Prior for Image Compression

Aug 2026 · IEEE Transactions on Image Processing · Vol 35, pp. 8966-8979 · 0 citations · 58 references
Computer Science Medicine

Abstract

Current neural image compression (NIC) methods primarily focus on signal fidelity optimization. While perceptually optimized codecs can generate decoded images that better align with human visual preferences at equivalent bitrates, they raise authenticity concerns due to potential deviations from the original content. Therefore, achieving controllable decoding is crucial in various applications. This study presents a novel plug-and-play framework that leverages large vision-language model (LVLM) priors to balance fidelity and perception for existing NICs. Our approach consists of two key components: a scalable Low-Rank Adaptation scheme to controllably enhance the semantics of initially decoded images, and a two-stage agent-assisted decoding strategy with vision-language priors utilization. Specifically, the first stage extracts textual semantic information from an LVLM using decoded images enhanced by flexible fidelity-perception decoding, while the second stage effectively integrates semantic priors from LVLMs, further mitigating decoding semantic uncertainty and achieving higher-quality decoding. Extensive experiments on multiple benchmark datasets demonstrate that our method enables off-the-shelf NICs to achieve flexible control between optimal perceptual quality and signal fidelity.

View source

Similar papers

Preprint Aug 2026

FLM: Frequency-Aware Language Models for Generative Image Compression

Generative models have significantly improved the performance ceiling of image lossy compression at low bitrates by exploiting learned priors. However, the generated textures and semantic details may deviate from the source content, thereby affecting the fidelity of image reconstruction. To solve these challenges, we propose FLM, a frequency-aware language model that improves compression efficiency through frequency-domain probabilistic modeling while retaining deterministic reconstruction. At the encoder, the input image is transformed into quantized DCT coefficients, which are organized into discrete sequences using macroblock-based coefficient tokenization. FLM then performs next-coefficient prediction to autoregressively estimate token-wise conditional probability distributions for arithmetic coding, thereby generating a compact bitstream. At the decoder, the LLM and arithmetic decoder jointly recover the frequency-domain data, followed by inverse transformations for image reconstruction. A task-specific frequency-domain dataset and a two-stage fine-tuning strategy are further developed to enable the model to operate across multiple bitrate settings. FLM is a versatile compressor that is compatible with both lossy compression and lossless JPEG recompression frameworks. Experiments show that FLM exceeds conventional and generative lossy compression methods in rate-distortion performance. FLM achieves BD-PSNR gains of 3.30 dB, 3.83 dB, and 3.80 dB than JPEG baseline on Kodak, Tecnick, and CLIC2020, respectively. Better qualitative quality of FLM can be achieved in improving semantically high fidelity and suppressing blocking artifacts. FLM is also validated to be applicable to the lossless recompression task with competitive performance.

Jia-Run Chen, Kejun Wu, Li Li et al. · 0 citations
Preprint Aug 2026

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

This work systematically study MoE designs for vision encoder scaling and finds that fine-grained MoE topologies yield substantial gains over both dense and standard MoE counterparts, and proposes an auxiliary-loss-free balancing variant for better expert utilization, and designs a specialized MoE kernel to mitigate inference latency overhead.

Bonan Zhang, Shiyu Dong, Quan Hung Tran et al. · 1 citation
Aug 2026

AdaMultiGAN: an adaptive multiscale decoding framework for few-shot image generation

An adaptive multi-scale decoding framework that effectively balances global context with fine-grained detail is proposed that exhibits superior robustness and generalization across diverse domains, effectively alleviating limitations of existing fusion-based approaches.

Yu Luo, Chunna Zhao, Yaqun Huang · 0 citations
Jul 2026

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

This work introduces \method, a Multi-scale Adaptive Vision Encoder, a Multi-scale Adaptive Vision Encoder that uses position-dependent gates to fuse shallow, intermediate, and deep features from a vision Transformer, preserving global semantics while enhancing edges, text, and local structure.

Sha Lei · 0 citations
Review Open access 2026

Semantic Image Compression for Rate, Distortion, and Performance Optimization: A Systematic Review with Quantitative Synthesis

The semantic image compression is revolutionizing the intelligent visual communication by compressing the image representation for machine perception and downstream tasks instead of human viewing. Although there have been great improvements with learned compression, semantic communication, and task-aware coding, there are still problems with architecture design, evaluation protocols, benchmark datasets, and deployment. This review, based on the PRISMA framework, examines 54 selected studies published from January 2018 to June 2026. Study-level coding was used to create a common coding scheme to perform a descriptive quantitative evidence synthesis, and frequencies and percentages calculated across learning architectures, optimization objectives, semantic preservation strategies, benchmark datasets, evaluation metrics, and application domains. The results demonstrate the straightforward shift from conventional rate–distortion optimization to semantic and task-aware compression. There are several major difficulties, including lack of standard evaluation frameworks, low cross-domain generalization, low reproducibility, and low unified metrics that combine compression efficiency and semantic fidelity. The review summarizes the existing methodologies, introduces new research trends, and offers useful insights for designing efficient semantic image compression systems for edge intelligence, medical imaging, autonomous platforms and communication networks empowered by artificial intelligence. The emphasis is on what steps can be taken towards learned image compression, which is a subject pursued by the Deep Learning (DL) community.

I. Manga, R. Mathew, B. Bali · 0 citations
Preprint Aug 2026

CodecArena: Codec Quality Assessment via Visual Reinforcement Learning

Video coding is advancing into the low and ultra-low bitrate regime, driven by end-to-end codecs that replace the hand-crafted pipeline with jointly optimized neural networks and generative codecs that exploit the priors of video generation models. Yet the dominant metrics, LPIPS and DISTS, measure feature and texture similarity rather than content fidelity: a reconstruction that hallucinates a wrong face or blurs text into convincing strokes can still score well, even when a human rejects it instantly. To address this, we propose CodecArena, the first vision-language framework for video coding quality assessment, casting codec evaluation as source-conditioned comparative reasoning between a reference and its reconstructions. We optimize CodecArena with Facet-GRPO, a visual reinforcement learning scheme that aligns pairwise codec preferences while grounding the verdict in five fidelity facets: identity, objects, text, texture, and temporal consistency. Its facet-anchored reward uses automatically derived facet directions as weak anchors, rather than human per-facet labels, to prevent any single sub-score from dominating the holistic preference and to yield interpretable fine-grained quality judgments. To support training and evaluation in this underexplored regime, we construct two complementary resources: CodecArena-1K, a fully automatic preference dataset of 1,500 comparison groups built from traditional, neural, and generative codec reconstructions with fused vision-language and objective supervision; and CodecArena-Bench, a human-ranked benchmark with source-disjoint videos for fair out-of-domain evaluation. Extensive experiments demonstrate that CodecArena achieves state-of-the-art agreement with human judgments on source-disjoint content across diverse codecs and bitrates, surpassing perceptual metrics and prior vision-language evaluators.

Jiaye Fu, Wei-Qi Li, Qiankun Gao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.