Aug 2026· ACM Transactions on Architecture and Code Optimization (TACO)· Vol 23, pp. 1-26· 0 citations· 15 references
TL;DR
NeuroComp is a cross-platform compiler that automatically generates functionally correct and high-performance code for heterogeneous neuromorphic backends, and introduces a unified multi-level loop abstraction to systematically align SNN computational semantics with hardware capacities.
Abstract
Spiking Neural Networks (SNNs) offer substantial energy efficiency advantages, yet deploying them across diverse neuromorphic hardware remains challenging. Existing toolchains exhibit a critical trade-off: platform-specific frameworks lack portability, while low-level programming interfaces demand extensive hardware expertise and produce non-portable code. Furthermore, conventional deep learning compilers fail to handle the unique temporal dynamics of SNNs and the idiosyncratic constraints of neuromorphic architectures. To bridge this gap, we present NeuroComp, a cross-platform compiler that automatically generates functionally correct and high-performance code for heterogeneous neuromorphic backends. Our core insight is that the diverse execution behaviors of neuromorphic hardware can be universally modeled as parallelizable scalar computations expressible via nested loops. Building on this, NeuroComp introduces a unified multi-level loop abstraction to systematically align SNN computational semantics with hardware capacities. To efficiently navigate the vast implementation space, we propose a decoupled two-stage scheduling strategy. The first stage performs batch partitioning to guarantee functional correctness and strict architectural compliance, while the second stage optimizes inter-batch execution orders for maximum data reuse, guided by lightweight, hardware-specific cost models and heuristic pruning. We evaluate NeuroComp on three representative architectural paradigms: an instruction-extended CPU, a dedicated accelerator, and a many-core hardware. Experimental results demonstrate that NeuroComp consistently achieves competitive performance compared to meticulously hand-optimized implementations and significantly outperforms existing generic automated compilers, effectively minimizing development effort while unlocking cross-platform portability.
The 194M-parameter model is implemented on an Alveo U50C FPGA using digital fixed-point arithmetic and on an ARM CPU using sparse integer execution to connect event sparsity to omitted computation and data movement in SymbolicLight V2.
A GPU-accelerated pipeline for SNN mapping is proposed: a multi-level partitioning scheme is devised around hardware constraints, while placement is initialized through recursive bisection, followed by refinement pulling together strongly connected cores through repeated swaps.
While Spiking Neural Networks offer a promising path toward energy-efficient edge intelligence, conventional Binary representations often suffer from high inference latency and discretization errors. Multi-level (graded) spike models address these limitations by increasing information density per pulse, yet they typi...
Wen-Fei Song, Andrea Castagnetti, Pierre-Emmanuel Novac et al.· Neuromorphic Computing and E...· 0 citations
This review surveys the full landscape of AI hardware accelerators for LLMs, including general-purpose GPUs, custom ASICs, reconfigurable FPGAs, processing-in-memory and near-memory architectures, and emerging neuromorphic and photonic approaches across cloud and edge deployment.
This work analyzes the Vitis AI compiler and proposes an XIR-level splitting framework that generates independently compilable .xmodel fragments while preserving the context required for DPU mapping, and restores correct DPU mapping by addressing boundary-context loss and incomplete dependency collection.
Federico Buccellato, Luca Mannini, C. De Sio· WiPiEC Journal - Works in Pr...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.