Aug 2026· WiPiEC Journal - Works in Progress in Embedded Computing Journal· 0 citations· 26 references
TL;DR
This work proposes a method to distribute the execution of individual layers across accelerators, and demonstrates how to implement such a baseline system using a SoC generator framework, performs an ablation study prototyping different versions on an FPGA, and identifies gaps and limitations by executing a multi-DNN autonomous driving application.
Abstract
Many machine learning applications require heterogeneous Deep Neural Networks (DNNs) to work collaboratively. Although several works have focused on how to serve these systems using cloud solutions, less attention has been paid to the edge scenarios. Particularly, coordinating heterogeneous AI workloads across a custom application-specific System-on-Chip (SoC) with multiple accelerators presents significant challenges. This work first demonstrates how to implement such a baseline system using a SoC generator framework, performs an ablation study prototyping different versions on an FPGA, details how an RTOS can be used to achieve model parallelism on multiple accelerators, and identifies gaps and limitations by executing a multi-DNN autonomous driving application. To improve the utilization of the system and increase throughput, we propose a method to distribute the execution of individual layers across accelerators. Instead of partitioning all layers in the same manner and statically allocating them at compile time, we select the ideal partitioning for each one during compilation using an autotuning process, and then dynamically assign them to the available accelerators during runtime. We analyze the variability in layer execution in a system with multiple accelerators and use this information to guide the runtime allocation of partitions. We demonstrate that our method achieves a mean 29 % and 40 % improvement in accelerator utilization and throughput over the model parallelism baseline, and a 10 % and 9 % improvement over a round-robin runtime distribution of partitions.
This thesis presents design methodologies that increase the extent to which key design choices are based on exploration and quantitative evaluation of design alternatives, and identifies architectures that consistently outperform the state-of-the-art, achieving considerable improvements in latency, throughput, energy,...
Procyon, a fine-grain multi-tenancy framework that fuses the PE instruction streams of multiple workloads into a unified execution schedule, substantially reduces PE underutilization that results in 3 × speedup over state-of-the-art sparse streaming accelerators, and reaches a peak throughput of 61 .
Ubaid Bakhtiar, Jeremy Sha, Helya Hosseini et al.· 0 citations
The Compatibility Ratio (CR) is introduced as a simple guideline for evaluating performance trade-offs between optimal hardware micro-architecture configurations across different workloads and shows that, for the considered accelerator, a DNN model-family optimized configuration might occupy an effective middle ground...
Lukas Groth, Andrija Nešković, Rainer Buchty et al.· ACM Transactions on Embedded...· 0 citations
With the continuous development of very large-scale integration (VLSI), the number of cores integrated on a single chip has reached hundreds. Network-on-Chip (NoC), featuring a highly scalable and high-bandwidth communication architecture, has been widely applied in Chip Multiprocessor Systems (CMP). NoC-based deep neu...
Cheng-Long Sun, Yi-He Zhang, Yajun Liu et al.· Journal of King Saud Univers...· 0 citations
. The fast developing pace in deep learning is giving pressure on computing hardware continuously. In many practical cases, model size and training cost increase faster than the performance improvement of general-purpose processors. The huge different pace between them makes a gap called “ compute gap ” . Consequently,...
Kang-Zhe Peng· Proceedings of the 3rd Inter...· 0 citations
Sparse DNN accelerators specialize in ANN or SNN execution, leaving energy or latency on the table when workload characteristics vary within a layer. Hybrid accelerator designs that switch modes at layer or tile granularity suffer from low PE utilization since one core type idles whenever the other is active. NeuroFlex...
Varun Manjunath, P. Ramesh, Gopalakrishnan Srinivasan· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.