Skip to content
Review Open access

Analysis of Research Progress on Deployment Methods for Deep Learning Models on FPGAs

Aug 2026 · Electronics · 0 citations · 55 references

TL;DR

A systematic review of FPGA-based DL deployment from a cross-layer perspective spanning model, compiler, architecture, runtime, and electronic design automation (EDA) is presented, highlighting that reliable cross-study comparison requires careful consideration of model configuration, precision, execution phase, batch size, memory residency, FPGA platform, and evidence maturity.

Abstract

Deep learning (DL) models have achieved remarkable progress in natural language processing, computer vision, content generation, and edge intelligence; however, their rapidly increasing computational complexity, memory demand, and deployment diversity pose significant challenges for practical implementation. Field-programmable gate arrays (FPGAs) provide customized low-precision computation, spatial dataflow, on-chip data reuse, reconfigurability, and rich I/O capabilities, making them an important platform for DL inference. This paper presents a systematic review of FPGA-based DL deployment from a cross-layer perspective spanning model, compiler, architecture, runtime, and electronic design automation (EDA). Following a PRISMA-guided evidence synthesis protocol, this review analyzes DL workload characteristics, FPGA architectural optimizations, deployment toolflows, and physical implementation challenges. A unified taxonomy is proposed along the specialization–programmability continuum, including model-fixed accelerators, generator-based accelerators, template-configurable accelerators, and ISA-programmable overlays. These approaches are compared according to hardware regeneration requirements, model adaptability, operator coverage, compilation cost, and deployment flexibility. Furthermore, emerging workloads, including vision Transformers, graph neural networks, large language models, and multimodal models, are analyzed from the perspectives of computation, memory behavior, and runtime coordination. The review shows that FPGA deployment efficiency increasingly depends on memory capacity, mutable state management, operator support, and end-to-end compilation capability rather than peak multiply–accumulate throughput alone. Based on the analysis of 70 primary FPGA implementation studies, this paper highlights that reliable cross-study comparison requires careful consideration of model configuration, precision, execution phase, batch size, memory residency, FPGA platform, and evidence maturity. For multimodal generative models, the current evidence remains limited, with no identified end-to-end FPGA-based vision–language model implementation in the reviewed corpus. This review provides a systematic perspective for future FPGA-based DL deployment research, emphasizing cross-layer optimization, physically aware compilation, extensible accelerator architectures, and practical deployment efficiency.

Read PDF

Similar papers

Review Aug 2026

Model Compression and Hardware-Aware Acceleration for Deep Learning on FPGAs: A Co-Design Taxonomy and Comparative Analysis

Deploying deep neural networks on Field-Programmable Gate Arrays (FPGAs) requires joint reasoning about model compression and hardware acceleration, however the most comprehensive existing cross-platform treatment of this space, Deng et al.~\cite{deng2020model}, compared compression techniques against CPU, GPU, FPGA, a...

Peter Forcha, H. Kajekusumadhar, Mbua Peter et al. · 0 citations
Conference Aug 2026

Designing and Building an FPGA Accelerator That Uses Less Energy for DNN Inference

Deep Neural Networks (DNNs) are critical to modern AI applications, yet their deployment on standard CPUs and GPUs is constrained by high power consumption and computational latency, particularly in resource-constrained edge environments. To address these limitations, this paper presents the design and implementation o...

P. V. G. K. Rao, Dudekula Raziya · 0 citations

Systematic Design Methodologies for Multi-Engine Deep Learning Accelerators

This thesis presents design methodologies that increase the extent to which key design choices are based on exploration and quantitative evaluation of design alternatives, and identifies architectures that consistently outperform the state-of-the-art, achieving considerable improvements in latency, throughput, energy,...

Fareed Mohammad Qararyah · 0 citations
Aug 2026

A Flexible Framework for Layer-Parallel CNN Training on FPGA Clusters

We present a flexible and scalable hardware/software framework for training CNNs on Ethernet-connected FPGA clusters using tightly pipelined layer parallelism. Starting from a high-level CNN and cluster description, the system automatically maps layers onto (potentially heterogeneous) devices, generates FPGA-specific b...

Philipp Kreowsky, Justin Knapheide, B. Stabernack · 1 citation
Conference Aug 2026

An Efficient HLS-Based Hardware Accelerator with Resource Optimization for Transformer Models

Deploying Transformer models on FPGA and System-on-Chip (SoC) platforms remains challenging due to their substantial computational complexity, large memory footprint, and high hardware resource requirements, particularly in multi-head attention and stacked encoder-decoder layers. This paper proposes a hardware-efficien...

Xuan Thao Tran, Thi Diem Tran · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.