Preprint
Aug 2026
Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs
This work proposes A-PACK, a two-stage framework that defers audio pruning until query-conditioned multimodal interactions emerge inside the LLM, and shows that audio exhibits higher task-relevant information density and representational diversity per token than video.
Kyeong-Jae Lee, Hongyeob Kim, Youngeun Kim et al.
· 2 citations