Skip to content
Preprint

DeVIT: Low-Power Vision Transformer Acceleration Using Delta Computation

Aug 2026 · 0 citations · 38 references
Computer Science

TL;DR

DeVIT, an acceleration method for vision transformers that leverages differential computation to enable multiplier-less matrix multiplication and introduces another useful property: value locality.

Abstract

The emergence of transformer-based deep learning models has brought unprecedented performance across various domains, particularly in natural language processing and computer vision. However, deploying these models, especially on resource-constrained devices, poses significant challenges due to their high computational complexity and large memory size and bandwidth requirements. This complexity has led researchers to use low-bit model weights to reduce memory usage and improve efficiency. In addition to reducing processing and memory demands, quantization introduces another useful property: value locality, where the extremely large number of parameters are restricted to a limited range of values. To fully take advantage of this locality, this paper presents DeVIT, an acceleration method for vision transformers that leverages differential computation to enable multiplier-less matrix multiplication.

View source

Similar papers

Preprint Sep 2026

PENDA: An Efficient Processing Element via Norm-of-Difference for Deep Learning Accelerators

Inner product computation dominates the computational cost of deep learning models; thus, accelerating this primitive is key to improving hardware efficiency. However, most existing techniques rely on approximations, which can degrade model accuracy. To preserve exactness while optimizing hardware, this paper presents...

Kai-Chieh Hsu, T. Chang · 0 citations
Conference Open access Sep 2026

SeGO: Sensitivity-Aware Golden Optimization for Large-Scale VLM Quantization

A cross-modal structural sensitivity asymmetry in VLMs is revealed and SeGO is proposed, a unified structural sensitivity-aware sparse optimization framework that achieves the balance among model parameter amount, quantization accuracy and scaling factors’ search efficiency on InternVL2 and LLaVA series.

Tian-Qi Zhao, Xin-Rui Cheng, Yang Su et al. · 0 citations
#edge computing Open access Sep 2026

Fine-Grained Energy Assessment of Vision Transformer Operators on Hybrid SoC Platforms: A CPU, iGPU, and NPU Case Study

As CPUs alone are often insufficiently efficient for running Vision Transformers (ViTs), contemporary computing platforms, particularly edge systems, increasingly adopt heterogeneous architectures that combine CPUs with GPUs and Neural Processing Units (NPUs). This paper presents a detailed energy assessment of individ...

Li-Xian Ho, C. Kachris · 0 citations
Preprint Aug 2026

APT: Accelerating Diffusion Transformers via Attention Probability-Guided Pruning and Quantization

APT, a software-hardware co-designed accelerator for high-resolution DiTs that leverages attention probabilities as a unified importance metric to jointly optimize computation through fine-grained pruning and adaptive precision scaling, and is evaluated on SOTA DiT models including PixArt, Stable Diffusion 3, and FLUX.

Sungyeob Yoo, Seeyeon Kim, Joonyong Park et al. · 0 citations
#edge computing Preprint Aug 2026

LiteEvent-AE: Lightweight Autoencoder for Event-Based Vision on Low-Latency Energy-Constrained Edge Devices

A compact and configurable event-driven autoencoder that efficiently compresses neuromorphic data while preserving essential spatiotemporal structure for downstream inference and demonstrates the potential of compact event-driven models to advance environmentally conscious, low-power AI systems for high-speed perceptio...

Riadul Islam, Joey Mulé, Dhandeep Challagundla et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.