Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· pp. 3621-3630· 0 citations· 21 references
TL;DR
MMSep is proposed, a training-free multimodal separator localization and compression framework that improves efficiency in both prefilling and decoding and consistently reduces latency while maintaining competitive generation quality and reasoning accuracy.
Abstract
Existing research on Efficient Multimodal Large Language Models (EMLLMs) primarily focuses on reducing the number of visual tokens in the prefilling stage, which is tailored to short-answer inference scenario. However, in more complex multimodal reasoning tasks, models are often required to generate lengthy intermediate reasoning rationales, while repeatedly revisiting prefilled contexts to verify and revise reasoning paths. As the generation length increases, the cumulative overhead of the decoding stage rises surpasses that of pruned prefilling, to become the dominant cost source for end-to-end inference. Investigating decoding-time attention behaviors, we observe two phenomena on textual and visual side related to selective and effective memory retention. Based on these observations, we propose MMSep, a training-free multimodal separator localization and compression framework that improves efficiency in both prefilling and decoding. MMSep (i) localizes visual anchors/separators during prefilling via question-guided attention and a spatial–similarity constraint, and (ii) performs structured KV-cache compression during decoding by retaining textual separators as long-range context and enabling separator-triggered, on-demand visual recall. Experiments on four MLLM backbones across long-generation and standard reasoning benchmarks demonstrate that MMSep consistently reduces latency while maintaining competitive generation quality and reasoning accuracy. Our code is available at https://github.com/MeinhardMark/MMSep.
Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units,...
Xu-Dong Wang, Hao Wu, Hao-Zhe Hu et al.· 0 citations
Modality-Contrastive Preference Optimization (MCPO), a highly sample-efficient two-stage length-compression method that requires fewer than 900 training samples, and adopts a highly nonlinear odds-ratio formulation that provides steep gradients in the with-image context to reinforce length constraints for preferred tra...
Guang-Heng Yang, Zhen-Liang Ni, Zhen-Kai Wu et al.· 0 citations
The inference efficiency of Multimodal Large Language Models (MLLMs) is severely constrained by massive visual token sequences induced by high-resolution inputs, with computational cost scaling quadratically. Existing approaches primarily focus on downstream token compression, while overlooking a fundamental upstream i...
Zhi-Qiang Xia, Yang Li, Xin-Yuan Zhang et al.· 0 citations
ResComEmb is proposed, a trainable framework for effective and efficient universal multi-vector multimodal embedding that produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval.
Zi-Jing Cai, Yu-Zhe Wang, Jing-Xian Zhu et al.· 0 citations
Multimodal Large Language Models (MLLMs), pre-trained on vast image-text datasets, excel in zero-shot reasoning but face catastrophic forgetting when adapting to sequential tasks in Continual Instruction Tuning (CIT). While lightweight prompt-based methods offer a parameter-efficient strategy for sequential adaptation,...
Hong-Sheng Zhang, Zhong Ji, Di Wang et al.· IEEE Transactions on Image P...· 0 citations
VLZip is introduced, a framework that unifies visual and textual compression for high-fidelity reasoning within a pure Transformer, and establishes an efficient and powerful new standard for long-context multimodal AI.
Yu-Qi Zhang, Cheng Chen, Yuyu Guo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.