A GPU-accelerated pipeline for SNN mapping is proposed: a multi-level partitioning scheme is devised around hardware constraints, while placement is initialized through recursive bisection, followed by refinement pulling together strongly connected cores through repeated swaps.
Abstract
SNNs running on neuromorphic hardware use spikes to achieve sparse and energy-efficient communication over a mesh of cores. In turn, system performance heavily depends on the assignment of neurons to cores: the mapping. Since hardware features inter-core multicast and intra-core replication of spikes, we model SNNs as hypergraphs to exploit both opportunities for reducing communication traffic. Mapping thus comprises two NP-hard problems: hypergraph partitioning and placement on the lattice of cores. High-quality solutions to both are critical, yet increasingly difficult as networks scale to millions of neurons. Therefore, we propose a GPU-accelerated pipeline for SNN mapping: a multi-level partitioning scheme is devised around hardware constraints, while placement is initialized through recursive bisection, followed by refinement pulling together strongly connected cores through repeated swaps. Model-based experiments show upwards of 16% lower latency and 42% lower energy for spike movements over existing sequential tools, while our parallel mapper is on average 18-280x faster.
A reconfigurable hybrid ring-shaped architecture (Rhr-NoC) to adapt to the large-scale data transmission patterns within accelerators and design a simulated annealing algorithm tailored to the hybrid ring-shaped architecture, providing a more rational scheme for mapping neural networks onto NoC platforms.
Cheng-Long Sun, Yi-He Zhang, Yajun Liu et al.· Journal of King Saud Univers...· 0 citations
Modern GPUs are no longer physically symmetric. Die scaling leads to both manufacturing-driven floorsweeping and cache and memory partitioning. The former creates chip-specific compute topologies, while the latter causes non-uniform memory access. These asymmetries are substantial. Topology-oblivious compute unit alloc...
Xiao-Ze Fan, Jian-Hao Wang, Wei-Hao Cui et al.· 0 citations
Mixture-of-experts (MoE) models expand capacity via sparse activation, but inference across GPUs introduces tensor-parallel (TP) collectives and expert-parallel (EP) dispatch and combine operations. Completion time depends not just on communication volume but on how logical groups map onto intra-server interconnects, G...
To mitigate interconnect scaling bottlenecks $\left(O\left(N^{2}\right)\right)$ and Non-Uniform Memory Access (NUMA) congestion in Programmable Multi-Core Accelerators (PMCAs), this paper introduces a multi-cluster architecture that replaces inter-cluster communication with localized data replication within ScratchPad...
Chanon Khongprasongsiri, P. Tanguy, Kevin J. M. Martin et al.· IEEE International Conferenc...· 0 citations
NeuroComp is a cross-platform compiler that automatically generates functionally correct and high-performance code for heterogeneous neuromorphic backends, and introduces a unified multi-level loop abstraction to systematically align SNN computational semantics with hardware capacities.
Chao Xiao, Lei Wang, Ya-Shuai Lu et al.· ACM Transactions on Architec...· 0 citations
This work compares two recent multicore neuromorphic systems implemented in the same 22-nm FDSOI technology and explicitly optimized for inter-core event communication, and discusses routing-aware training as a means of jointly optimizing neural connectivity, task performance, and hardware mappability.
Sebastian Billaudelle, Christian Metzner, Jimmy Weber et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.