Skip to content
Preprint

PENDA: An Efficient Processing Element via Norm-of-Difference for Deep Learning Accelerators

Sep 2026 · 0 citations · 18 references
Computer Science

Abstract

Inner product computation dominates the computational cost of deep learning models; thus, accelerating this primitive is key to improving hardware efficiency. However, most existing techniques rely on approximations, which can degrade model accuracy. To preserve exactness while optimizing hardware, this paper presents PENDA (processing element via norm-of-difference architecture), which leverages the law of cosines to recast multiplications as squared-difference operations. Replacing multiply-accumulate units with the proposed norm-of-difference units yields 11~36%, 5~48%, and 11~19% reductions in area, energy, and clock period, respectively, for the PE array of a deep learning accelerator.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.