Skip to content
Conference

Memory-Aware Architectural Exploration Method to Design Programmable Multi-Core Accelerators

Sep 2026 · IEEE International Conference on Application-Specific Systems, Architectures, and Processors · pp. 169-170 · 0 citations · 4 references

Abstract

To mitigate interconnect scaling bottlenecks $\left(O\left(N^{2}\right)\right)$ and Non-Uniform Memory Access (NUMA) congestion in Programmable Multi-Core Accelerators (PMCAs), this paper introduces a multi-cluster architecture that replaces inter-cluster communication with localized data replication within ScratchPad Memories (SPMs). This design creates a critical architectural trade-off between the data replication overhead required by overlapping-footprint kernels (e.g., matrix multiplication) and the routing delays of larger clusters. We propose an architectural exploration methodology utilizing a Memory-Delay Squared Product $(m d^{2} p)$ cost function to analytically resolve this trade-off. Linear programming optimization reveals two optimal configurations based on the workload data ratio: a 1-PE cluster for element-wise kernels and an 8-PE cluster for overlapping-footprint kernels. Synthesized results for a 64-PE implementation demonstrate that our methodology yields substantial performance gains, achieving up to a $7.36 \times$ speedup over an ARM Cortex A53 and a $6.03 \times$ speedup over a state-of-the-art integrated programmable array (IPA) accelerator.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.