This work presents an HLS tool that embeds fine-grained pipeline control in a sequential programming model, enabling rapid design-space exploration and outperforms HLS tools with sequential semantics and achieve PPA comparable to hand-written RTL.
Abstract
High-level synthesis (HLS) raises the abstraction of hardware design from concurrent register-transfer level (RTL) programs to sequential programs. Among the forms of parallelism HLS exploits, pipelining demands fine-grained control over pipeline structure and hazard resolution to achieve competitive power, performance, and area (PPA). However, existing tools either lack such control or sacrifice sequential semantics to provide it. We present an HLS tool that embeds fine-grained pipeline control in a sequential programming model, enabling rapid design-space exploration. The tool builds on visibility control, a novel programming abstraction that unifies hazard resolution strategies including stalling, bypassing, speculation, deferred commit, and register renaming. We evaluate on in-order RISC-V cores, histograms, and an AES accelerator. On RISC-V cores, we implement stall, bypass, speculation, and register renaming; on histograms, we implement scheduling strategies that previously required RTL or concurrent programming models. Compiled pipelines outperform HLS tools with sequential semantics and achieve PPA comparable to hand-written RTL.
This paper proposes a compile-time instruction scheduling method that models sub-cycle resource usage and analyzes both data and structural dependencies at fine granularity, enabling precise detection of structural hazards in complex execution units.
Dolunay is introduced, a RISC-V-based Independent Thread Scheduling (ITS) SIMT accelerator that employs a cooperative multitasking model and explicit synchronization barriers at the hardware-level, and provides the forward-progress guarantees necessary to implement starvation-free algorithms.
Ahmet Can, Erkan Uslu· WiPiEC Journal - Works in Pr...· 0 citations
Exo-GPU, an imperative, low-level language that creates minimal abstraction over CUDA, is proposed, to treat parallelism and synchronization as mere annotations on sequential code rather than as fundamental control flow primitives, enabling verification that these constructs do not alter the program semantics.
David Akeley, Yuka Ikarashi, Jonathan Ragan-Kelley· 0 citations
Loom, a tuning-free symbolic compiler framework for tile-based SPMD programs on spatial dataflow architectures, is presented, suggesting that hardware-derived symbolic compilation provides a retargetable alternative to profiling-based tuning for spatial dataflow architectures while remaining interpretable by keeping op...
He-Ru Wang, Wei Li, Zhen-Yu Bai et al.· 0 citations
Dataflow-based coarse-grained reconfigurable architectures (CGRAs) and dynamic high-level synthesis (DHLS) are both promising for accelerating applications with nontrivial control and memory behavior, but existing compilation flows are typically fragmented and often struggle with control handling, memory ordering, and...
Yi Huang, Xiang-Yu Kong, Jian-Feng Zhu et al.· IEEE Transactions on Paralle...· 0 citations
This paper presents Ouros, a pipelined processor for lazy functional programming languages based on combinator graph reduction. Ouros overcomes the inherent sequentiality of graph reduction through dataflow-driven execution and automatic fine-grained multi-threading. It maintains pipeline utilisation by hiding per-thre...
Yu-Kang Xie, Craig R. Ramsay, Robert J. Stewart et al.· IEEE International Conferenc...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.