Skip to content
Open access

NeuroComp: A Cross-Platform Compiler for Efficient SNN Deployment on Neuromorphic Hardware

Aug 2026 · ACM Transactions on Architecture and Code Optimization (TACO) · Vol 23, pp. 1-26 · 0 citations · 15 references

TL;DR

NeuroComp is a cross-platform compiler that automatically generates functionally correct and high-performance code for heterogeneous neuromorphic backends, and introduces a unified multi-level loop abstraction to systematically align SNN computational semantics with hardware capacities.

Abstract

Spiking Neural Networks (SNNs) offer substantial energy efficiency advantages, yet deploying them across diverse neuromorphic hardware remains challenging. Existing toolchains exhibit a critical trade-off: platform-specific frameworks lack portability, while low-level programming interfaces demand extensive hardware expertise and produce non-portable code. Furthermore, conventional deep learning compilers fail to handle the unique temporal dynamics of SNNs and the idiosyncratic constraints of neuromorphic architectures. To bridge this gap, we present NeuroComp, a cross-platform compiler that automatically generates functionally correct and high-performance code for heterogeneous neuromorphic backends. Our core insight is that the diverse execution behaviors of neuromorphic hardware can be universally modeled as parallelizable scalar computations expressible via nested loops. Building on this, NeuroComp introduces a unified multi-level loop abstraction to systematically align SNN computational semantics with hardware capacities. To efficiently navigate the vast implementation space, we propose a decoupled two-stage scheduling strategy. The first stage performs batch partitioning to guarantee functional correctness and strict architectural compliance, while the second stage optimizes inter-batch execution orders for maximum data reuse, guided by lightweight, hardware-specific cost models and heuristic pruning. We evaluate NeuroComp on three representative architectural paradigms: an instruction-extended CPU, a dedicated accelerator, and a many-core hardware. Experimental results demonstrate that NeuroComp consistently achieves competitive performance compared to meticulously hand-optimized implementations and significantly outperforms existing generic automated compilers, effectively minimizing development effort while unlocking cross-platform portability.

Read PDF

Similar papers

Preprint Sep 2026

GPU-Accelerated Hypergraph Partitioning and Placement to Map SNNs on Neuromorphic Hardware

A GPU-accelerated pipeline for SNN mapping is proposed: a multi-level partitioning scheme is devised around hardware constraints, while placement is initialized through recursive bisection, followed by refinement pulling together strongly connected cores through repeated swaps.

Marco Ronzani, C. Silvano · 0 citations
Open access Oct 2026

Beyond binary SNNs: a hardware-algorithm co-design and architectural evaluation for graded spike networks

While Spiking Neural Networks offer a promising path toward energy-efficient edge intelligence, conventional Binary representations often suffer from high inference latency and discretization errors. Multi-level (graded) spike models address these limitations by increasing information density per pulse, yet they typi...

Wen-Fei Song, Andrea Castagnetti, Pierre-Emmanuel Novac et al. · 0 citations
Review Aug 2026

AI Hardware Accelerators for Large Language Models: Architectures and the Memory Wall

This review surveys the full landscape of AI hardware accelerators for LLMs, including general-purpose GPUs, custom ASICs, reconfigurable FPGAs, processing-in-memory and near-memory architectures, and emerging neuromorphic and photonic approaches across cloud and edge deployment.

Siddharth Patel, Rohit Singh · 0 citations
Open access Aug 2026

A Compiler-Aware Framework for Partitioned Neural Network Inference on FPGA DPUs

This work analyzes the Vitis AI compiler and proposes an XIR-level splitting framework that generates independently compilable .xmodel fragments while preserving the context required for DPU mapping, and restores correct DPU mapping by addressing boundary-context loss and incomplete dependency collection.

Federico Buccellato, Luca Mannini, C. De Sio · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.