Skip to content

ATLAS: Automated HLS for DL-Optimized FPGAs

Jul 2026 · arXiv.org · Vol abs/2607.07643 · 0 citations · 17 references
Computer Science

TL;DR

ATLAS is presented, a fully automated flow from a high-level DL model description to a hardware implementation on an FPGA with custom in-fabric DL-optimized hardblocks, requiring no manual RTL design or explicit hardblock instantiation from the end user.

Abstract

FPGA architectures increasingly incorporate domain-specific in-fabric hardblocks to accelerate DL inference, particularly GEMM, which dominates DL computation. To realize the performance gains of these hardblocks, manual RTL design is required: the programmer must understand the hardblock microarchitecture, instantiate them in RTL, and manage tiling and control logic. While programming in C/C++ and using HLS tools has increased the abstraction level and productivity of FPGA engineers, HLS tools do not support code generation for custom hardblocks natively. Prior work has demonstrated that blackbox mechanisms in HLS tools can be used to target custom hardblocks, but this still requires explicit function calls in user-written HLS C and manual creation of RTL IP libraries, significant effort that must be repeated for every layer in a DL model. Furthermore, for DL, an even high-level programming interface, e.g., Pytorch/Keras instead of C/C++, is desirable for improved programmability and user adoption. We present ATLAS, a fully automated flow from a high-level DL model description to a hardware implementation on an FPGA with custom in-fabric DL-optimized hardblocks, requiring no manual RTL design or explicit hardblock instantiation from the end user. Our approach uses GEMM as a universal abstraction layer and comprises two components: (1) hls4ml-GEMM, a compiler frontend that transforms DL layers into HLS C code with architecture-agnostic GEMM function calls, and (2) a GEMM IP Generator, an architecture-aware backend that produces hardblock-based RTL wrappers with tiling logic, control FSMs, and scheduling metadata. We evaluate the flow across 11 DL designs, including individual fully connected, convolution, and attention layers, as well as full CNN, MLP, and Transformer models targeting an FPGA architecture with Tensor Slices using Catapult for HLS and VTR for implementation.

View source

Similar papers

Jul 2026

High-Level Synthesis of Efficient Pipelines with Visibility Control

This work presents an HLS tool that embeds fine-grained pipeline control in a sequential programming model, enabling rapid design-space exploration and outperforms HLS tools with sequential semantics and achieve PPA comparable to hand-written RTL.

Jungin Rhee, Minseong Jang, Jaewoo Kim et al. · 0 citations
Preprint Aug 2026

HLSmith: An Expert-Guided Agentic Framework for C/C++-to-HLS Translation

HLSmith, an expert-guided framework for translating C/C++ programs into optimized HLS accelerators, is presented and evaluated on PolyBench against ChatHLS, a leading prior agent-orchestration framework for HLS accelerator development.

Yuebo Luo, Ahmad Sedigh Baroughi, Philip Stachura et al. · 0 citations
Jul 2026

Pattern-Guided Design Space Exploration for FPGA Accelerator Design

PATTERNDSE is presented, a lightweight pattern-guided design space exploration (DSE) framework for FPGA kernels written in Allo, a scheduling-oriented HLS programming system, demonstrating that computation-pattern information can prune unproductive schedule combinations while preserving high-quality HLS outcomes.

Jialiang Zhang, Weiman Yan, Yue-Lin Zou · 0 citations

Automated Generation of RISC-V Extensions with Formal Correctness Guarantees

Janus is presented, an LLM-assisted framework that synthe-sizes custom instructions integrated into the Ibex RISC-V core while keeping correctness outside the agent, demonstrating a practical path for using LLMs to explore ISA specialization without making the agent part of the trusted correctness boundary.

Elisavet Lydia Alvanaki, Jia-Kun Wang, Eugenio Muscinelli et al. · 0 citations
Book Open access Aug 2026

POSTER: VibeNIC: Toward LLM-driven Agile Development of FPGA SmartNICs

VibeNIC is a SmartNIC framework co-designed at every layer for an LLM developer, and a case study on a stateful HBM-augmented UDP datapath delivers a working end-to-end design in hours.

Yunfan Li, Jialin Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.