Skip to content
Book Open access

MMSep: Efficient Multimodal Long-Generation Reasoning via Multimodal Separator Compression

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 3621-3630 · 0 citations · 21 references

TL;DR

MMSep is proposed, a training-free multimodal separator localization and compression framework that improves efficiency in both prefilling and decoding and consistently reduces latency while maintaining competitive generation quality and reasoning accuracy.

Abstract

Existing research on Efficient Multimodal Large Language Models (EMLLMs) primarily focuses on reducing the number of visual tokens in the prefilling stage, which is tailored to short-answer inference scenario. However, in more complex multimodal reasoning tasks, models are often required to generate lengthy intermediate reasoning rationales, while repeatedly revisiting prefilled contexts to verify and revise reasoning paths. As the generation length increases, the cumulative overhead of the decoding stage rises surpasses that of pruned prefilling, to become the dominant cost source for end-to-end inference. Investigating decoding-time attention behaviors, we observe two phenomena on textual and visual side related to selective and effective memory retention. Based on these observations, we propose MMSep, a training-free multimodal separator localization and compression framework that improves efficiency in both prefilling and decoding. MMSep (i) localizes visual anchors/separators during prefilling via question-guided attention and a spatial–similarity constraint, and (ii) performs structured KV-cache compression during decoding by retaining textual separators as long-range context and enabling separator-triggered, on-demand visual recall. Experiments on four MLLM backbones across long-generation and standard reasoning benchmarks demonstrate that MMSep consistently reduces latency while maintaining competitive generation quality and reasoning accuracy. Our code is available at https://github.com/MeinhardMark/MMSep.

Read PDF

Similar papers

#artificial intelligence Preprint Oct 2026

MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs

Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units,...

Xu-Dong Wang, Hao Wu, Hao-Zhe Hu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression

Modality-Contrastive Preference Optimization (MCPO), a highly sample-efficient two-stage length-compression method that requires fewer than 900 training samples, and adopts a highly nonlinear odds-ratio formulation that provides steep gradients in the with-image context to reinforce length constraints for preferred tra...

Guang-Heng Yang, Zhen-Liang Ni, Zhen-Kai Wu et al. · 0 citations
Preprint Sep 2026

Resolution as a First-Class Decision: Task-Conditioned Routing for Efficient Multimodal Large Language Models

The inference efficiency of Multimodal Large Language Models (MLLMs) is severely constrained by massive visual token sequences induced by high-resolution inputs, with computational cost scaling quadratically. Existing approaches primarily focus on downstream token compression, while overlooking a fundamental upstream i...

Zhi-Qiang Xia, Yang Li, Xin-Yuan Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression

ResComEmb is proposed, a trainable framework for effective and efficient universal multi-vector multimodal embedding that produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval.

Zi-Jing Cai, Yu-Zhe Wang, Jing-Xian Zhu et al. · 0 citations
Sep 2026

Interleaved Prompt Generation for Continual Instruction Tuning of Multimodal Large Language Model

Multimodal Large Language Models (MLLMs), pre-trained on vast image-text datasets, excel in zero-shot reasoning but face catastrophic forgetting when adapting to sequential tasks in Continual Instruction Tuning (CIT). While lightweight prompt-based methods offer a parameter-efficient strategy for sequential adaptation,...

Hong-Sheng Zhang, Zhong Ji, Di Wang et al. · 0 citations
Preprint Aug 2026

VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling

VLZip is introduced, a framework that unifies visual and textual compression for high-fidelity reasoning within a pure Transformer, and establishes an efficient and powerful new standard for long-context multimodal AI.

Yu-Qi Zhang, Cheng Chen, Yuyu Guo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.