Skip to content

CodeShrink: Adaptive Visual Compression for Efficient Multimodal Code Understanding

Jul 2026 · arXiv.org · Vol abs/2607.29637 · 1 citation · 41 references
Computer Science

TL;DR

Results show that combining layout compaction, adaptive configuration, and instruction-aware pruning can make multimodal code understanding more efficient, and propose CodeShrink, an adaptive visual compression framework with three components.

Abstract

Rendering source code as images offers a promising way to reduce the input costs of Multimodal Large Language Models (MLLMs). Adjusting image resolution can trade visual token cost against content fidelity. However, resolution scaling alone overlooks two sources of inefficiency: blank regions created by line breaks and indentation, and code regions irrelevant to the current instruction. Moreover, the best compression setting varies across inputs, tasks, and models, limiting fixed-ratio strategies. We propose CodeShrink, an adaptive visual compression framework with three components. Blank-Free Rendering replaces whitespace-dependent layouts with compact layouts and explicit structural markers, removing layout-induced tokens. Adaptive Compression Configuration uses a lightweight agent trained with reinforcement learning to predict a per-input setting that balances token efficiency and readability. Dominant Token Selection jointly analyzes the instruction and code image to prune task-irrelevant visual tokens during inference. We evaluate CodeShrink on code question answering, clone detection, and code completion. CodeShrink reduces visual token use by up to 71.2\% while matching or exceeding uncompressed text-only inputs, and consistently outperforms text-based and visual compression baselines across all three tasks. These results show that combining layout compaction, adaptive configuration, and instruction-aware pruning can make multimodal code understanding more efficient. Our code is available at https://github.com/vinsontang1/CodeShrink.

View source

Similar papers

Preprint Aug 2026

Instruction Distillation: Text Instructions as Visual Examples

Instruction Distillation is proposed: an offline procedure in which the MLLM itself generates, for each individual training image, a structured identification instruction encoding general appearance cues, features that differentiate the class from visually similar ones, and a common confusion point.

Hardik Jindal, Soumyabrata Pal, Sayak Ray Chowdhury · 0 citations
Preprint Aug 2026

SEER: Long-Context Reasoning via Selective Visual-Text Compression

SEER is presented, a framework that learns to select query-relevant images through visual scanning and retrieve textual content only where needed, combining the efficiency of visual compression with the precision of text-based reasoning.

Jiawei Xu, Zhilin Zhai, Jinrui Fang et al. · 3 citations
#computer vision Preprint Sep 2026

On the Design Fundamentals of Pixel Text Representation Learning

This work investigates the fundamental design principles required for robust visual text representation learning and trains Pixel Linguist II, a native-resolution vision encoder trained with on-the-fly rendering, unified contrastive grounding, and a multilingual curriculum over 280M training examples.

Chaohao Yuan, Rui-Feng Yuan, Zhuoxu Huang et al. · 1 citation
Preprint Aug 2026

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression

SPIRAL (Self-improving Path Integration and Realignment), a self-supervised alignment framework that closes this gap using only the model's own text-path behavior as supervision, requiring no external teachers or additional annotations, generalize to out-of-domain benchmarks, confirming that effective VTC hinges on ali...

Tianyu Liang, Xiangxi Zheng, Yilin Wang et al. · 2 citations
Preprint Aug 2026

InstructionCrafter: Generating Consistent and High-Fidelity Visual Instructions

Given textual task instructions, generating step-by-step visual instructions as an image sequence requires the simultaneous satisfaction of multiple properties, specifically step faithfulness, cross-image consistency, and per-frame visual quality. Existing text-to-image generation approaches rarely meet all three prope...

S. Okamoto, Satoshi Iizuka, Kazuhiro Fukui · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.