Skip to content

Compatibility Ratio as a Guideline for Hardware/DNN Co-Design of Embedded Accelerators

Sep 2026 · ACM Transactions on Embedded Computing Systems · 0 citations · 10 references
Advanced Neural Network Applications

TL;DR

The Compatibility Ratio (CR) is introduced as a simple guideline for evaluating performance trade-offs between optimal hardware micro-architecture configurations across different workloads and shows that, for the considered accelerator, a DNN model-family optimized configuration might occupy an effective middle ground between highly targeted single- and general-purpose configurations.

Abstract

Increasingly large deep neural networks (DNNs) pose a significant challenge regarding required compute capabilities and energy consumption, especially at the edge. This challenge generally necessitates dedicated inference hardware accelerators. Hardware/software co-design can improve overall system performance by jointly optimizing the hardware micro-architecture and DNN workload mapping. It is however not easy to quantify how well such an accelerator generalizes to workloads not considered during hardware/software co-design. In response to this, we introduce the Compatibility Ratio (CR) as a simple guideline for evaluating performance trade-offs between optimal hardware micro-architecture configurations across different workloads. CR allows us to quantify the performance trade-offs of deploying a workload on an accelerator optimized for a different workload. We demonstrate CR through two case studies on a systolic array-based accelerator. First, we apply CR to explore the design space of the accelerator across 13 DNN workloads. In this case study, CR analysis showed that the choice of representative workload during co-design can implicitly increase the normalized area-latency cost of unconsidered workloads by more than 30% in the evaluated design space. Furthermore, we use CR to analyze how well our systolic array-based accelerator template can generalize beyond a single DNN workload to cover a family of DNN workloads. Our findings show that, for the considered accelerator, a DNN model-family optimized configuration might occupy an effective middle ground between highly targeted single- and general-purpose configurations. Second, we use CR as a guideline for a practical memory-retargeting decision in a specialized variant of our accelerator template. In this case study, CR quantifies whether a memory-retargeted accelerator derivative is justified under the selected memory-area and latency objective. For this two-configuration memory-retargeting case, analytical CR differs from implementation-level CR by 0.01, corresponding to one percentage point on the normalized CR scale.

View source

Similar papers

Conference Aug 2026

Designing and Building an FPGA Accelerator That Uses Less Energy for DNN Inference

Deep Neural Networks (DNNs) are critical to modern AI applications, yet their deployment on standard CPUs and GPUs is constrained by high power consumption and computational latency, particularly in resource-constrained edge environments. To address these limitations, this paper presents the design and implementation o...

P. V. G. K. Rao, Dudekula Raziya · 0 citations
Preprint Aug 2026

Evaluating MFU as a Proxy for GPU Power for Energy-Aware Simulation of LLM Training

This work evaluates whether Model FLOPs Utilization (MFU) can serve as a portable, software-defined predictor of GPU power for LLMs, finding that a linear MFU-based power model fits every tested GPU as long as the workload is compute-bound, as in production LLM training.

Niklas Enskat, Philipp Wiesner · 1 citation · ⚡1
Open access Aug 2026

Autotuned Distribution of Multi-DNN Workloads on Multi-Accelerator SoCs

This work proposes a method to distribute the execution of individual layers across accelerators, and demonstrates how to implement such a baseline system using a SoC generator framework, performs an ablation study prototyping different versions on an FPGA, and identifies gaps and limitations by executing a multi-DNN a...

Federico Nicolás Peccia, Avik Bhatnagar, Oliver Bringmann · 0 citations
Preprint Sep 2026

AccelForge: Comprehensive Modeling and Co-Design Framework for AI Accelerators

Tensor algebra workloads, of which deep neural networks are prominent examples, are energy-intensive workloads in modern datacenter and edge deployments, making accelerators necessary to achieve energy efficiency and high throughput. To quickly evaluate and iterate on accelerator designs, we need an accelerator modelin...

Tanner Andrulis, Michael Gilbert, Vivienne Sze et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 29, 2026

Introducing Quine: An AI research system designed for the complexity of biology

Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.