Skip to content

EVA: Efficient Token Compression Co-Optimization With Heterogeneous CIM Architecture for Video LLM Acceleration

Sep 2026 · IEEE Transactions on Very Large Scale Integration (vlsi) Systems · Vol 34, pp. 2935-2948 · 0 citations · 52 references

Abstract

Video large language models (VLLMs) enable powerful multimodal reasoning but face severe efficiency challenges on edge devices due to the massive computational and memory demands caused by lengthy video-token sequences. Existing methods struggle to efficiently compress spatiotemporally redundant tokens while minimizing DRAM access overhead during inference. In this work, we propose EVA, a co-designed algorithm–hardware framework for accelerating VLLMs. At the algorithm level, we introduce an efficient training-free token compression (ETC) method that combines greedy temporal segmentation (GTS) to adaptively partition frames by content similarity, with diversity spatiotemporal compression (DSC) to retain semantically rich tokens from both static and dynamic regions. The method is plug-and-play, requires no retraining, and employs sign similarity to enable hardware-friendly computing at scale. At the hardware level, we design a heterogeneous accelerator integrating a lightweight token compression engine (TCE), a computing-in-memory (CIM) engine for in situ execution of linear layers, and a reconfigurable digital attention engine with hardware-specialized attention computation and an interleaved pipeline dataflow. Across multiple models and video benchmarks, EVA preserves accuracy under aggressive token reduction and delivers substantial efficiency gains. Notably, on LLaVA-OneVision-7B, EVA compresses 90% of video tokens while maintaining 97.8% of the original accuracy. EVA achieves up to <inline-formula> <tex-math notation="LaTeX">$10.3\times $ </tex-math></inline-formula> speedup and <inline-formula> <tex-math notation="LaTeX">$47.4\times $ </tex-math></inline-formula> energy reduction compared to GPU, and achieves up to <inline-formula> <tex-math notation="LaTeX">$3.8\times $ </tex-math></inline-formula> speedup over prior specialized accelerators, demonstrating a practical path toward scalable VLLM inference on resource-constrained devices.

View source

Similar papers

Book Open access Aug 2026

MVP: A Mobile 3D-Stacked VLM Accelerator for Efficient Video Understanding by Leveraging Dynamic Sparse Attention Patterns

MVP, a 3D-stacked VLM accelerator featuring context-aware sparse attention (CASA) and online workload-aware hybrid parallelism scheduling, is introduced, which prunes redundant attention computation FLOPS and adaptively balances computation across hybrid bonding (HB) based many-core NoC architecture.

Yifan Ding, Qianxu Wang, Dunshan Yu et al. · 0 citations
Open access Sep 2026

Benchmarking token mixers for efficient transformer-based learned video compression

Learned video compression has rapidly evolved, with recent approaches demonstrating potential in complex context modeling. However, these performance gains often come at the cost of significant complexity in both framework design and computation-heavy modules, obscuring the efficiency contribution of the core architect...

Chun Zhang, He-Ming Sun, J. Katto · 0 citations
#artificial intelligence Preprint Aug 2026

Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding

Faster Flash Decoding (FFD) is presented, a novel hardware-algorithm co-design framework designed to break the memory wall in long-context decoding and introduces the top-delta strategy, which dynamically filters blocks to achieve distribution-adaptive sparsity without global synchronization.

Zhigeng Liu, Zhiyuan Ning, Rui-Xiao Li et al. · 3 citations
Open access Sep 2026

DCSR-GCN: A High-Performance GCN Accelerator Based on Dynamic Compression and Sparsity Reordering

Graph Convolutional Networks (GCNs) are widely used in tasks involving irregular graph data, such as recommendation. The hybrid execution pattern of sparse aggregation and dense combination during inference limits the efficiency of general processors like CPU and GPU. Therefore, designing dedicated accelerators for GCN...

Jun-Sheng Chang, Yi-Min Zhao, Yu-Xin Huang et al. · 0 citations
Preprint Aug 2026

Streamable Neural Video Compression: A Mixed Precision Approach for Cross-Platform Deployment

Neural Video Codecs (NVCs) offer unprecedented rate-distortion performance, making them highly attractive for bandwidth-constrained environments like 5G cellular networks and emerging satellite direct-to-cell (D2C) links. However, deploying NVCs in real-world streaming applications is severely hindered by cross-platfor...

Kasidis Arunruangsirilert, He-Ming Sun, J. Katto · 0 citations

Related blog posts

Microsoft Research Blog Sep 29, 2026

Introducing Quine: An AI research system designed for the complexity of biology

Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.