Skip to content
Conference Open access

A Multi-Dimensional Evaluation Framework for Matrix-Based AI Accelerators: GPU, FPGA, and ASIC

2026 · Proceedings of the 3rd International Conference on Mechanics, Electronics Engineering and Automation, ICMEEA 2026, April 24-26, 2026, Singapore, Singapore · 0 citations

Abstract

. The fast developing pace in deep learning is giving pressure on computing hardware continuously. In many practical cases, model size and training cost increase faster than the performance improvement of general-purpose processors. The huge different pace between them makes a gap called “ compute gap ” . Consequently, heterogeneous accelerators such as GPUs, FPGAs, and ASICs have become central to modern AI systems. It is not a easy work to select the right and appropriate accelerator because current studies focus on peak throughput while ignore the latency characteristics, energy efficiency and deployment constraints. This essay introduces a multi-dimensional evaluation framework for matrix-based AI accelerators which concentrated on convolution-dominated workload. Analyzing the represented hardware platform, including NVIDIA A100 GPUs, AMD Xilinx Versal AI Core FPGAs and Google TPU v4 ASICs, and comparing their throughput, determinism, scalability and power consumption. At the same time, in order to provide the guide in selection of hardware in different and detailed scenarios and distinguish compute-bound and memory-bound workloads, the Roofline Model is utilized. The analysis suggests that GPUs and TPUs are generally more suitable for large-scale cloud training, whereas FPGAs tend to be more advantageous in latency-sensitive and energy-constrained edge inference scenarios.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.