PENDA: An Efficient Processing Element via Norm-of-Difference for Deep Learning Accelerators
Inner product computation dominates the computational cost of deep learning models; thus, accelerating this primitive is key to improving hardware efficiency. However, most existing techniques rely on approximations, which can degrade model accuracy. To preserve exactness while optimizing hardware, this paper presents...