Memory-Aware Architectural Exploration Method to Design Programmable Multi-Core Accelerators
Abstract
To mitigate interconnect scaling bottlenecks $\left(O\left(N^{2}\right)\right)$ and Non-Uniform Memory Access (NUMA) congestion in Programmable Multi-Core Accelerators (PMCAs), this paper introduces a multi-cluster architecture that replaces inter-cluster communication with localized data replication within ScratchPad Memories (SPMs). This design creates a critical architectural trade-off between the data replication overhead required by overlapping-footprint kernels (e.g., matrix multiplication) and the routing delays of larger clusters. We propose an architectural exploration methodology utilizing a Memory-Delay Squared Product $(m d^{2} p)$ cost function to analytically resolve this trade-off. Linear programming optimization reveals two optimal configurations based on the workload data ratio: a 1-PE cluster for element-wise kernels and an 8-PE cluster for overlapping-footprint kernels. Synthesized results for a 64-PE implementation demonstrate that our methodology yields substantial performance gains, achieving up to a $7.36 \times$ speedup over an ARM Cortex A53 and a $6.03 \times$ speedup over a state-of-the-art integrated programmable array (IPA) accelerator.