Skip to content
Preprint

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs

Aug 2026 · 2 citations · 32 references
Computer Science

TL;DR

This work proposes A-PACK, a two-stage framework that defers audio pruning until query-conditioned multimodal interactions emerge inside the LLM, and shows that audio exhibits higher task-relevant information density and representational diversity per token than video.

Abstract

Omni-modal LLMs jointly process audio, video, and text, but long multimodal sequences incur substantial prefill and KV-cache costs. Existing omni-modal compression methods primarily focus on pre-LLM token reduction, leaving modality-specific compression across the LLM boundary underexplored. We propose A-PACK, a two-stage framework that defers audio pruning until query-conditioned multimodal interactions emerge. Our analysis shows that audio exhibits higher task-relevant information density and representational diversity per token than video. We further find that local audio-visual dynamics provide a more effective cue for visual selection than token-wise matching. We therefore preserve audio and compress video with local dynamics before the LLM, then progressively prune low-relevance audio and visual tokens and their KV-cache entries inside the LLM. Across four benchmarks on Qwen2.5-Omni-7B/3B, A-PACK achieves the strongest average performance among the evaluated prior methods while reducing prefill FLOPs by up to 78% and improving decoding throughput by up to 2.21x.

View source

Similar papers

Preprint Aug 2026

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio token sequences incurs substantial computational overhead, demanding aggressive token compression for efficient deployment. Existing methods of...

Wan-Shun Su, Yang Shi, Fei Liu et al. · 2 citations
Preprint Sep 2026

OmniRoute: Mapping Temporal Semantic Evidence to Audio-Visual Token Budgets for Efficient Omnimodal Large Language Models

Omnimodal large language models (Omni-LLMs) encode audio and visual streams into temporally interleaved token sequences for multimodal reasoning. However, processing long audio-visual token sequences incurs substantial prefill costs. Existing compression methods have made progress, but often overlook temporal changes i...

Yu-Chen Deng, Zi-Dang Cai, Fei-Diao Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception

Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence witho...

Ju-Yi Lin, Zhi-Qiang Lao, Jia-Li Cui et al. · 0 citations
#computer vision Review Sep 2026

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal grounding comes at a co...

Killian Steunou, Yannis Tevissen, M. E. El Yacoubi · 0 citations
Sep 2026

AdaCompVL: Adaptive Compression of Spatiotemporal and Cross-Modal Redundancy for Efficient Video-Language Learning.

Multimodal large language models (MLLMs) have recently extended from static image understanding to video comprehension, but representing videos as frame-level token sequences incurs substantial computational overhead. Existing visual token compression methods typically rely on uniform sampling or single-dimension redun...

Jian-Xin Ma, Shi-Bo Jin, Lu-Juan Dang et al. · 0 citations
Preprint Sep 2026

CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding

Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and answer user questions under tight latency constraints. Existing methods improve efficiency through token pruning and memory-bank schemes, but mainly reduce visual tokens after visual encoding. Consequently, downst...

Jing-Chi Jiang, Yi-Ran Ling, Ruo-Nan Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.