Skip to content

High-Level Synthesis of Efficient Pipelines with Visibility Control

Jul 2026 · arXiv.org · Vol abs/2607.18765 · 0 citations · 50 references
Computer Science

TL;DR

This work presents an HLS tool that embeds fine-grained pipeline control in a sequential programming model, enabling rapid design-space exploration and outperforms HLS tools with sequential semantics and achieve PPA comparable to hand-written RTL.

Abstract

High-level synthesis (HLS) raises the abstraction of hardware design from concurrent register-transfer level (RTL) programs to sequential programs. Among the forms of parallelism HLS exploits, pipelining demands fine-grained control over pipeline structure and hazard resolution to achieve competitive power, performance, and area (PPA). However, existing tools either lack such control or sacrifice sequential semantics to provide it. We present an HLS tool that embeds fine-grained pipeline control in a sequential programming model, enabling rapid design-space exploration. The tool builds on visibility control, a novel programming abstraction that unifies hazard resolution strategies including stalling, bypassing, speculation, deferred commit, and register renaming. We evaluate on in-order RISC-V cores, histograms, and an AES accelerator. On RISC-V cores, we implement stall, bypass, speculation, and register renaming; on histograms, we implement scheduling strategies that previously required RTL or concurrent programming models. Compiled pipelines outperform HLS tools with sequential semantics and achieve PPA comparable to hand-written RTL.

View source

Similar papers

Open access Aug 2026

Fine-Grained Structural Conflict Modeling for Compile-Time Instruction Scheduling on VLIW ASIPs

This paper proposes a compile-time instruction scheduling method that models sub-cycle resource usage and analyzes both data and structural dependencies at fine granularity, enabling precise detection of structural hazards in complex execution units.

Peng Hao, Shengbing Zhang, Xinbing Zhou et al. · 0 citations
Open access Aug 2026

Dolunay: Architectural Support for Independent Thread Scheduling in a RISC-V SIMT Accelerator

Dolunay is introduced, a RISC-V-based Independent Thread Scheduling (ITS) SIMT accelerator that employs a cooperative multitasking model and explicit synchronization barriers at the hardware-level, and provides the forward-progress guarantees necessary to implement starvation-free algorithms.

Ahmet Can, Erkan Uslu · 0 citations
Preprint Sep 2026

Exo-GPU: Safe, Imperative, User-schedulable Programming for Tensor Cores

Exo-GPU, an imperative, low-level language that creates minimal abstraction over CUDA, is proposed, to treat parallelism and synchronization as mere annotations on sequential code rather than as fundamental control flow primitives, enabling verification that these constructs do not alter the program semantics.

David Akeley, Yuka Ikarashi, Jonathan Ragan-Kelley · 0 citations
Preprint Sep 2026

Schedules Are Solvable Symbols: Tuning-Free Compilation of Tile Programs on Dataflow Architectures

Loom, a tuning-free symbolic compiler framework for tile-based SPMD programs on spatial dataflow architectures, is presented, suggesting that hardware-derived symbolic compilation provides a retargetable alternative to profiling-based tuning for spatial dataflow architectures while remaining interpretable by keeping op...

He-Ru Wang, Wei Li, Zhen-Yu Bai et al. · 0 citations
Nov 2026

RedPanda: A Unified Compilation Framework for Dataflow-Based CGRAs and High-Level Synthesis

Dataflow-based coarse-grained reconfigurable architectures (CGRAs) and dynamic high-level synthesis (DHLS) are both promising for accelerating applications with nontrivial control and memory behavior, but existing compilation flows are typically fragmented and often struggle with control handling, memory ordering, and...

Yi Huang, Xiang-Yu Kong, Jian-Feng Zhu et al. · 0 citations
Conference Sep 2026

Ouros: A Dataflow-Driven Processor for Lazy Functional Programming Languages

This paper presents Ouros, a pipelined processor for lazy functional programming languages based on combinator graph reduction. Ouros overcomes the inherent sequentiality of graph reduction through dataflow-driven execution and automatic fine-grained multi-threading. It maintains pipeline utilisation by hiding per-thre...

Yu-Kang Xie, Craig R. Ramsay, Robert J. Stewart et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.