Sep 2026· IEEE Transactions on Very Large Scale Integration (vlsi) Systems· Vol 34, pp. 2935-2948· 0 citations· 52 references
Abstract
Video large language models (VLLMs) enable powerful multimodal reasoning but face severe efficiency challenges on edge devices due to the massive computational and memory demands caused by lengthy video-token sequences. Existing methods struggle to efficiently compress spatiotemporally redundant tokens while minimizing DRAM access overhead during inference. In this work, we propose EVA, a co-designed algorithm–hardware framework for accelerating VLLMs. At the algorithm level, we introduce an efficient training-free token compression (ETC) method that combines greedy temporal segmentation (GTS) to adaptively partition frames by content similarity, with diversity spatiotemporal compression (DSC) to retain semantically rich tokens from both static and dynamic regions. The method is plug-and-play, requires no retraining, and employs sign similarity to enable hardware-friendly computing at scale. At the hardware level, we design a heterogeneous accelerator integrating a lightweight token compression engine (TCE), a computing-in-memory (CIM) engine for in situ execution of linear layers, and a reconfigurable digital attention engine with hardware-specialized attention computation and an interleaved pipeline dataflow. Across multiple models and video benchmarks, EVA preserves accuracy under aggressive token reduction and delivers substantial efficiency gains. Notably, on LLaVA-OneVision-7B, EVA compresses 90% of video tokens while maintaining 97.8% of the original accuracy. EVA achieves up to <inline-formula> <tex-math notation="LaTeX">$10.3\times $ </tex-math></inline-formula> speedup and <inline-formula> <tex-math notation="LaTeX">$47.4\times $ </tex-math></inline-formula> energy reduction compared to GPU, and achieves up to <inline-formula> <tex-math notation="LaTeX">$3.8\times $ </tex-math></inline-formula> speedup over prior specialized accelerators, demonstrating a practical path toward scalable VLLM inference on resource-constrained devices.
EdgeXpert is proposed, a software-hardware co-designed LLM accelerator that resolves this incompatibility and achieves up to 56.3% latency reduction and 44.1% energy reduction compared to prior works, while maintaining near-baseline accuracy.
Sangwoo Ha, Hyunwoo Seo, Yurim Jo et al.· 0 citations
Learned video compression has rapidly evolved, with recent approaches demonstrating potential in complex context modeling. However, these performance gains often come at the cost of significant complexity in both framework design and computation-heavy modules, obscuring the efficiency contribution of the core architect...
Chun Zhang, He-Ming Sun, J. Katto· APSIPA Transactions on Signa...· 0 citations
Faster Flash Decoding (FFD) is presented, a novel hardware-algorithm co-design framework designed to break the memory wall in long-context decoding and introduces the top-delta strategy, which dynamically filters blocks to achieve distribution-adaptive sparsity without global synchronization.
Zhigeng Liu, Zhiyuan Ning, Rui-Xiao Li et al.· 3 citations
Graph Convolutional Networks (GCNs) are widely used in tasks involving irregular graph data, such as recommendation. The hybrid execution pattern of sparse aggregation and dense combination during inference limits the efficiency of general processors like CPU and GPU. Therefore, designing dedicated accelerators for GCN...
Jun-Sheng Chang, Yi-Min Zhao, Yu-Xin Huang et al.· ACM Transactions on Design A...· 0 citations
Neural Video Codecs (NVCs) offer unprecedented rate-distortion performance, making them highly attractive for bandwidth-constrained environments like 5G cellular networks and emerging satellite direct-to-cell (D2C) links. However, deploying NVCs in real-world streaming applications is severely hindered by cross-platfor...
Kasidis Arunruangsirilert, He-Ming Sun, J. Katto· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…