Sep 2026· IEEE International Conference on Application-Specific Systems, Architectures, and Processors· pp. 85-92· 0 citations· 15 references
Abstract
This paper presents Ouros, a pipelined processor for lazy functional programming languages based on combinator graph reduction. Ouros overcomes the inherent sequentiality of graph reduction through dataflow-driven execution and automatic fine-grained multi-threading. It maintains pipeline utilisation by hiding per-thread latency and interleaving multiple independent threads. Its concurrent garbage collector (GC) addresses the high memory allocation pressure of functional language execution. The correctness of the GC algorithm is verified via model checking. Ouros achieves a higher clock frequency than the single-cycle KappaMutor processor in FPGA implementation, and is faster by 20.9% on average across 10 Haskell benchmarks (up to 118% on richly-threaded programs). GC overhead ranges from 0% to 23% depending on program allocation behaviour.
Dolunay is introduced, a RISC-V-based Independent Thread Scheduling (ITS) SIMT accelerator that employs a cooperative multitasking model and explicit synchronization barriers at the hardware-level, and provides the forward-progress guarantees necessary to implement starvation-free algorithms.
Ahmet Can, Erkan Uslu· WiPiEC Journal - Works in Pr...· 0 citations
Exo-GPU, an imperative, low-level language that creates minimal abstraction over CUDA, is proposed, to treat parallelism and synchronization as mere annotations on sequential code rather than as fundamental control flow primitives, enabling verification that these constructs do not alter the program semantics.
David Akeley, Yuka Ikarashi, Jonathan Ragan-Kelley· 0 citations
Loom, a tuning-free symbolic compiler framework for tile-based SPMD programs on spatial dataflow architectures, is presented, suggesting that hardware-derived symbolic compilation provides a retargetable alternative to profiling-based tuning for spatial dataflow architectures while remaining interpretable by keeping op...
He-Ru Wang, Wei Li, Zhen-Yu Bai et al.· 0 citations
Interpreters have a large indirect-branch footprint, requiring large predictor capacity for accurate prediction. We propose a hardware/software co-design in which a hardware lookahead engine, running ahead of the pipeline with software-provided bytecode metadata, supplies interpreter dispatch targets to the frontend. T...
The conventional threading model multiplexes software threads onto hardware cores. This model inherently suffers from the overhead of 1) context switching, 2) scheduling, and 3) system event notifications (e.g., I/O interrupts). As computing enters μs-scale, such overheads become the key bottleneck of a wide range of d...
Yi-Ming Yao, Xiao-He Qin, Yi Fan et al.· Proceedings of the ACM SIGOP...· 0 citations
The increasing diversity of parallel hardware challenges existing compilation flows. While OpenMP provides a portable abstraction for shared-memory parallelism, existing compilers tightly couple the frontend semantics with fixed lowering strategies. This design limits performance portability across different runtimes a...
Luca Parigi, Giuseppe Tagliavini· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.