Skip to content

Systematic Design Methodologies for Multi-Engine Deep Learning Accelerators

TL;DR

This thesis presents design methodologies that increase the extent to which key design choices are based on exploration and quantitative evaluation of design alternatives, and identifies architectures that consistently outperform the state-of-the-art, achieving considerable improvements in latency, throughput, energy, and energy-delay product (EDP).

Abstract

Domain-Specific Accelerators (DSAs) have become a key driver of performance and efficiency improvements in the post-Moore's Law era. These improvements stem from specializing the hardware for domain workloads and exploiting the workloads' inherent parallelism. In the domain of Deep Learning (DL), workloads (models) comprise multiple operations, known as layers, that exhibit parallelism opportunities and diverse computational characteristics. Consequently, DSAs with multiple computational units (engines) provide a natural architectural paradigm to fully exploit the specialization and parallelism potential inherent in such multi-layered models.Multi-engine DL accelerators generally fall into two categories: model-specific and flexible. Model-specific accelerators are co-designed to efficiently execute one or a few closely related models. Flexible accelerators, by contrast, are designed to support a broad range of DL workloads. Designing and implementing accelerators in either category that fully exploit specialization and parallelism, and thus optimize performance and efficiency, requires systematic exploration based on quantitative evaluation of design alternatives. Existing multi-engine DL accelerator design approaches range from intuition-driven to exploration-based methodologies. However, even the latter typically leave key architectural parameters unexplored by fixing them a priori based on expert knowledge and intuition. In many cases, the fixed parameters are more consequential for accelerator specialization and parallelism than the explored parameters. Consequently, the full potential of the multi-engine paradigm is often left unexploited.To fully exploit the potential of multi-engine DL accelerators, this thesis presents design methodologies that increase the extent to which key design choices are based on exploration and quantitative evaluation of design alternatives. The first contribution of this thesis is the Fixed Budget Hybrid CNN Accelerator (FiBHA). FiBHA proposes a hybrid, model-specific, multi-engine architecture and an accompanying design methodology. FiBHA targets a specific class of DL models and relies primarily on empirical analysis to exploit opportunities for specialization and parallelism. To expand the scope and co-design accelerators for a broader class of models, the work moved to a more analytical approach. The second contribution of this thesis comprises MCCM, a fast analytical cost model for evaluating model-specific multi-engine accelerators, and MCExplorer, a design space exploration framework built upon it. Together, they enable orders-of-magnitude faster evaluation and systematic exploration of model-specific multi-engine accelerator designs. Unlike existing approaches that rely on predefined design choices, MCExplorer quantitatively evaluates alternative architectural configurations across a broader design space. To further expand the scope, the work extends to flexible, in addition to model-specific, multi-engine accelerators. The third contribution of the thesis is a design methodology for flexible multi-engine DL accelerators, termed MEDEM. To support a wide range of diverse DL workloads, a flexible multi-engine accelerator must have an engine combination with complementary capabilities to ensure that different layers across these diverse workloads are processed efficiently. Existing work builds flexible multi-engine accelerators by combining expert-selected, independently optimized engines. However, independently optimized engines may perform best on largely overlapping subsets of workloads, and thus their combination does not necessarily improve overall workload coverage. MEDEM presents an alternative design methodology where the engines are co-designed, then curated to find a combination that maximizes the coverage of diverse workloads.Using a systematic approach based on modeling and quantitative evaluation of a wider space of design alternatives, the proposed methodologies identify accelerator architectures that better exploit the specialization and parallelism inherent in DL workloads. This applies to both model-specific accelerators, as FiBHA, MCCM, and MCExplorer demonstrate, and to flexible ones, as MEDEM shows. As a result, these methodologies identify architectures that consistently outperform the state-of-the-art, achieving considerable improvements in latency, throughput, energy, and energy-delay product (EDP).

View source

Similar papers

Open access Aug 2026

Autotuned Distribution of Multi-DNN Workloads on Multi-Accelerator SoCs

This work proposes a method to distribute the execution of individual layers across accelerators, and demonstrates how to implement such a baseline system using a SoC generator framework, performs an ablation study prototyping different versions on an FPGA, and identifies gaps and limitations by executing a multi-DNN a...

Federico Nicolás Peccia, Avik Bhatnagar, Oliver Bringmann · 0 citations
Preprint Sep 2026

AccelForge: Comprehensive Modeling and Co-Design Framework for AI Accelerators

Tensor algebra workloads, of which deep neural networks are prominent examples, are energy-intensive workloads in modern datacenter and edge deployments, making accelerators necessary to achieve energy efficiency and high throughput. To quickly evaluate and iterate on accelerator designs, we need an accelerator modelin...

Tanner Andrulis, Michael Gilbert, Vivienne Sze et al. · 0 citations

Procyon: Promoting Fine-Grain Multi-Tenancy to Optimize Sparse Streaming Accelerators

Procyon, a fine-grain multi-tenancy framework that fuses the PE instruction streams of multiple workloads into a unified execution schedule, substantially reduces PE underutilization that results in 3 × speedup over state-of-the-art sparse streaming accelerators, and reaches a peak throughput of 61 .

Ubaid Bakhtiar, Jeremy Sha, Helya Hosseini et al. · 0 citations
Review Open access Aug 2026

Analysis of Research Progress on Deployment Methods for Deep Learning Models on FPGAs

A systematic review of FPGA-based DL deployment from a cross-layer perspective spanning model, compiler, architecture, runtime, and electronic design automation (EDA) is presented, highlighting that reliable cross-study comparison requires careful consideration of model configuration, precision, execution phase, batch...

Shuo Wang, Lei Chen, Chunsheng Tian et al. · 0 citations
Review Aug 2026

AI Hardware Accelerators for Large Language Models: Architectures and the Memory Wall

Large language models (LLMs) place unprecedented and still-growing demands on the hardware that trains and serves them. This review surveys the full landscape of AI hardware accelerators for LLMs, including general-purpose GPUs, custom ASICs such as TPUs, Trainium, Groq, and Cerebras, reconfigurable FPGAs, processing-i...

Siddharth Patel, Rohit Singh · 0 citations
Preprint Jul 2026

Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators

This work presents the first production-scale application of Triton on a custom ML accelerator, MTIA-2i, developed by Meta, and develops a new compiler backend that targets it, introduces enhancements to TorchInductor code generation, and proposes minimal language extensions that expose MTIA-specific architectural feat...

Haishan Zhu, Domi Yan, Michael Levesque-Dion et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.