Skip to content
Preprint

GPU-Accelerated Hypergraph Partitioning and Placement to Map SNNs on Neuromorphic Hardware

Sep 2026 · 0 citations · 39 references
Computer Science

TL;DR

A GPU-accelerated pipeline for SNN mapping is proposed: a multi-level partitioning scheme is devised around hardware constraints, while placement is initialized through recursive bisection, followed by refinement pulling together strongly connected cores through repeated swaps.

Abstract

SNNs running on neuromorphic hardware use spikes to achieve sparse and energy-efficient communication over a mesh of cores. In turn, system performance heavily depends on the assignment of neurons to cores: the mapping. Since hardware features inter-core multicast and intra-core replication of spikes, we model SNNs as hypergraphs to exploit both opportunities for reducing communication traffic. Mapping thus comprises two NP-hard problems: hypergraph partitioning and placement on the lattice of cores. High-quality solutions to both are critical, yet increasingly difficult as networks scale to millions of neurons. Therefore, we propose a GPU-accelerated pipeline for SNN mapping: a multi-level partitioning scheme is devised around hardware constraints, while placement is initialized through recursive bisection, followed by refinement pulling together strongly connected cores through repeated swaps. Model-based experiments show upwards of 16% lower latency and 42% lower energy for spike movements over existing sequential tools, while our parallel mapper is on average 18-280x faster.

View source

Similar papers

Open access Aug 2026

A novel simulated annealing based mapping for hybrid NoC-enabled DNN accelerators

A reconfigurable hybrid ring-shaped architecture (Rhr-NoC) to adapt to the large-scale data transmission patterns within accelerators and design a simulated annealing algorithm tailored to the hybrid ring-shaped architecture, providing a more rational scheme for mapping neural networks onto NoC platforms.

Cheng-Long Sun, Yi-He Zhang, Yajun Liu et al. · 0 citations
Preprint Sep 2026

Dissecting How Die Scaling Breaks GPU Fine-grained Scheduling

Modern GPUs are no longer physically symmetric. Die scaling leads to both manufacturing-driven floorsweeping and cache and memory partitioning. The former creates chip-specific compute topologies, while the latter causes non-uniform memory access. These asymmetries are substantial. Topology-oblivious compute unit alloc...

Xiao-Ze Fan, Jian-Hao Wang, Wei-Hao Cui et al. · 0 citations
Preprint Sep 2026

Joint Effects of GPU Server Topology, Parallelism, and Congestion Control on MoE Inference: A Controlled Simulation Study

Mixture-of-experts (MoE) models expand capacity via sparse activation, but inference across GPUs introduces tensor-parallel (TP) collectives and expert-parallel (EP) dispatch and combine operations. Completion time depends not just on communication volume but on how logical groups map onto intra-server interconnects, G...

Kai-Kai Yuan, Rui Xi, Yu Liu · 0 citations
Conference Open access Sep 2026

Memory-Aware Architectural Exploration Method to Design Programmable Multi-Core Accelerators

To mitigate interconnect scaling bottlenecks $\left(O\left(N^{2}\right)\right)$ and Non-Uniform Memory Access (NUMA) congestion in Programmable Multi-Core Accelerators (PMCAs), this paper introduces a multi-cluster architecture that replaces inter-cluster communication with localized data replication within ScratchPad...

Chanon Khongprasongsiri, P. Tanguy, Kevin J. M. Martin et al. · 0 citations
Open access Aug 2026

NeuroComp: A Cross-Platform Compiler for Efficient SNN Deployment on Neuromorphic Hardware

NeuroComp is a cross-platform compiler that automatically generates functionally correct and high-performance code for heterogeneous neuromorphic backends, and introduces a unified multi-level loop abstraction to systematically align SNN computational semantics with hardware capacities.

Chao Xiao, Lei Wang, Ya-Shuai Lu et al. · 0 citations
Preprint Aug 2026

Small-World Communication Fabrics for Neuromorphic Multicore-SoCs

This work compares two recent multicore neuromorphic systems implemented in the same 22-nm FDSOI technology and explicitly optimized for inter-core event communication, and discusses routing-aware training as a means of jointly optimizing neural connectivity, task performance, and hardware mappability.

Sebastian Billaudelle, Christian Metzner, Jimmy Weber et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.