Skip to content
Book Open access

MMSep: Efficient Multimodal Long-Generation Reasoning via Multimodal Separator Compression

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 3621-3630 · 0 citations · 3 references

Abstract

Existing research on Efficient Multimodal Large Language Models (EMLLMs) primarily focuses on reducing the number of visual tokens in the prefilling stage, which is tailored to short-answer inference scenario. However, in more complex multimodal reasoning tasks, models are often required to generate lengthy intermediate reasoning rationales, while repeatedly revisiting prefilled contexts to verify and revise reasoning paths. As the generation length increases, the cumulative overhead of the decoding stage rises surpasses that of pruned prefilling, to become the dominant cost source for end-to-end inference. Investigating decoding-time attention behaviors, we observe two phenomena on textual and visual side related to selective and effective memory retention. Based on these observations, we propose MMSep, a training-free multimodal separator localization and compression framework that improves efficiency in both prefilling and decoding. MMSep (i) localizes visual anchors/separators during prefilling via question-guided attention and a spatial–similarity constraint, and (ii) performs structured KV-cache compression during decoding by retaining textual separators as long-range context and enabling separator-triggered, on-demand visual recall. Experiments on four MLLM backbones across long-generation and standard reasoning benchmarks demonstrate that MMSep consistently reduces latency while maintaining competitive generation quality and reasoning accuracy. Our code is available at https://github.com/MeinhardMark/MMSep.

Read PDF

Similar papers

Jul 2026

SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models

It is observed that attention scores from both vision and text tokens peak at modality separator tokens, suggesting that these separators bridge the two modalities and proposes SepPrune, an efficient, training-free, plug-and-play pruning method that uses the separator token as a unified query to rank and select informa...

Yucheng Wang, Qihui Zhu, Yang Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression

Modality-Contrastive Preference Optimization (MCPO), a highly sample-efficient two-stage length-compression method that requires fewer than 900 training samples, and adopts a highly nonlinear odds-ratio formulation that provides steep gradients in the with-image context to reinforce length constraints for preferred tra...

Guang-Heng Yang, Zhen-Liang Ni, Zhen-Kai Wu et al. · 0 citations
Sep 2026

Interleaved Prompt Generation for Continual Instruction Tuning of Multimodal Large Language Model

Multimodal Large Language Models (MLLMs), pre-trained on vast image-text datasets, excel in zero-shot reasoning but face catastrophic forgetting when adapting to sequential tasks in Continual Instruction Tuning (CIT). While lightweight prompt-based methods offer a parameter-efficient strategy for sequential adaptation,...

Hong-Sheng Zhang, Zhong Ji, Di Wang et al. · 0 citations
Preprint Aug 2026

VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling

VLZip is introduced, a framework that unifies visual and textual compression for high-fidelity reasoning within a pure Transformer, and establishes an efficient and powerful new standard for long-context multimodal AI.

Yu-Qi Zhang, Cheng Chen, Yuyu Guo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Current solutions predominantly rely on auxiliary models for token reduction but face a fundamental dilemma: lightweight enco...

Hao-Yu Guo, Yuan Feng, Junlin Lv et al. · 1 citation
Book Open access Aug 2026

UniDocVLM: Enhancing Visual Reasoning and Document Understanding for VLM via Reinforcement Learning

UnicDocVLM is proposed, an end-to-end framework that unifies OCR and visual RAG within a single vision-language model: the model first generates a structured parse of retrieved pages, then activates question-relevant evidence from the parse to support grounded reasoning and answering.

Zong-Sheng Cao, Anran Liu, Jun Xie et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.