Skip to content
Open access

Autotuned Distribution of Multi-DNN Workloads on Multi-Accelerator SoCs

Aug 2026 · WiPiEC Journal - Works in Progress in Embedded Computing Journal · 0 citations · 26 references

TL;DR

This work proposes a method to distribute the execution of individual layers across accelerators, and demonstrates how to implement such a baseline system using a SoC generator framework, performs an ablation study prototyping different versions on an FPGA, and identifies gaps and limitations by executing a multi-DNN autonomous driving application.

Abstract

Many machine learning applications require heterogeneous Deep Neural Networks (DNNs) to work collaboratively. Although several works have focused on how to serve these systems using cloud solutions, less attention has been paid to the edge scenarios. Particularly, coordinating heterogeneous AI workloads across a custom application-specific System-on-Chip (SoC) with multiple accelerators presents significant challenges. This work first demonstrates how to implement such a baseline system using a SoC generator framework, performs an ablation study prototyping different versions on an FPGA, details how an RTOS can be used to achieve model parallelism on multiple accelerators, and identifies gaps and limitations by executing a multi-DNN autonomous driving application. To improve the utilization of the system and increase throughput, we propose a method to distribute the execution of individual layers across accelerators. Instead of partitioning all layers in the same manner and statically allocating them at compile time, we select the ideal partitioning for each one during compilation using an autotuning process, and then dynamically assign them to the available accelerators during runtime. We analyze the variability in layer execution in a system with multiple accelerators and use this information to guide the runtime allocation of partitions. We demonstrate that our method achieves a mean 29 % and 40 % improvement in accelerator utilization and throughput over the model parallelism baseline, and a 10 % and 9 % improvement over a round-robin runtime distribution of partitions.

Read PDF

Similar papers

Systematic Design Methodologies for Multi-Engine Deep Learning Accelerators

This thesis presents design methodologies that increase the extent to which key design choices are based on exploration and quantitative evaluation of design alternatives, and identifies architectures that consistently outperform the state-of-the-art, achieving considerable improvements in latency, throughput, energy,...

Fareed Mohammad Qararyah · 0 citations

Procyon: Promoting Fine-Grain Multi-Tenancy to Optimize Sparse Streaming Accelerators

Procyon, a fine-grain multi-tenancy framework that fuses the PE instruction streams of multiple workloads into a unified execution schedule, substantially reduces PE underutilization that results in 3 × speedup over state-of-the-art sparse streaming accelerators, and reaches a peak throughput of 61 .

Ubaid Bakhtiar, Jeremy Sha, Helya Hosseini et al. · 0 citations
#edge computing Sep 2026

Compatibility Ratio as a Guideline for Hardware/DNN Co-Design of Embedded Accelerators

The Compatibility Ratio (CR) is introduced as a simple guideline for evaluating performance trade-offs between optimal hardware micro-architecture configurations across different workloads and shows that, for the considered accelerator, a DNN model-family optimized configuration might occupy an effective middle ground...

Lukas Groth, Andrija Nešković, Rainer Buchty et al. · 0 citations
Open access Aug 2026

A novel simulated annealing based mapping for hybrid NoC-enabled DNN accelerators

With the continuous development of very large-scale integration (VLSI), the number of cores integrated on a single chip has reached hundreds. Network-on-Chip (NoC), featuring a highly scalable and high-bandwidth communication architecture, has been widely applied in Chip Multiprocessor Systems (CMP). NoC-based deep neu...

Cheng-Long Sun, Yi-He Zhang, Yajun Liu et al. · 0 citations
Conference Open access 2026

A Multi-Dimensional Evaluation Framework for Matrix-Based AI Accelerators: GPU, FPGA, and ASIC

. The fast developing pace in deep learning is giving pressure on computing hardware continuously. In many practical cases, model size and training cost increase faster than the performance improvement of general-purpose processors. The huge different pace between them makes a gap called “ compute gap ” . Consequently,...

Kang-Zhe Peng · 0 citations
Preprint Nov 2025

NeuroFlex: Lossless Element-Level ANN-SNN Co-Execution for Efficient Sparse Inference

Sparse DNN accelerators specialize in ANN or SNN execution, leaving energy or latency on the table when workload characteristics vary within a layer. Hybrid accelerator designs that switch modes at layer or tile granularity suffer from low PE utilization since one core type idles whenever the other is active. NeuroFlex...

Varun Manjunath, P. Ramesh, Gopalakrishnan Srinivasan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.