Jul 2026· IEEE International Symposium on High-Performance Parallel Distributed Computing· pp. 620-627· 0 citations· 28 references
Computer Science
TL;DR
This paper presents the first systematic study showing that memory contention, not network bandwidth, is the dominant bottleneck in message aggregation runtimes, and advocates for dedicated near-memory subsystems to improve the scalability and performance of message aggregation runtimes.
Abstract
Small-message aggregation is critical for scaling irregular, communication intensive applications in high-performance computing. In this paper, contrary to conventional wisdom, we present the first systematic study showing that memory contention, not network bandwidth, is the dominant bottleneck in message aggregation runtimes. Using the state-of-the-art conveyors library as our reference implementation, we conducted extensive experiments on HPC systems featuring Slingshot 11 and InfiniBand interconnects, scaling to 16k cores (256 nodes) and processing 10s–100s GB of data. Our measurements reveal that interference between user data and aggregation buffers drives LLC miss rates to 77%, inflating memory costs by 2–3× over the algorithmic baseline. Consequently, we advocate for dedicated near-memory subsystems to improve the scalability and performance of message aggregation runtimes. This paper also demonstrates up to an order of magnitude higher latency for conveyor termination compared to a traditional HPC barrier, and it examines the impact of communication context isolation and the critical challenge of programmability.
This study provides a comprehensive analysis of cache memory, including its historical evolution, hierarchical levels (L1–L3), architectural design, and functional significance in modern computing systems, and indicates that cache size alone does not present a statistically significant difference between AMD and Intel...
Bukie, P. T., Ogar, A. A., Obono, I. Ofem et al.· Global Journal of Pure and A...· 0 citations
Building on NCCL's device-side API, low-latency interfaces for constructing custom collective kernels are developed and used to implement new symmetric collectives in NCCL, demonstrating benefits for both AI inference and traditional HPC workloads.
Siyuan Shen, Anton Korzh, J. Bachan et al.· arXiv.org· 0 citations
This work proposes SAI, a mechanism that virtualizes shared memory into the L2 cache to improve GPU performance for AI applications and introduces an L2 cache management strategy that integrates associativity-based virtual page allocation and a replacement information table, reducing page-swapping overhead while preser...
Hanqing Li, Tie-Jun Li, Sheng Ma et al.· ACM Transactions on Design A...· 0 citations
The main result concerns the plain-load path: attained LDG bandwidth peaks at a small offered per-thread load (K ~ 2) and then declines, by about 35% from K=2 to K=8 at the authors' primary configuration.
The evolution of supercomputer architecture has undergone several transformative phases
since the 1960s, yet existing surveys have not adequately captured the engineering trade-offs
that defined each generation. This paper presents a critical survey of supercomputer
architecture from early vector machines to contemp...
Oraye Godspower· International Journal of Com...· 0 citations
MEPOWER is proposed, a flexible, model-based approach to exposing compute/data movement imbalance that characterizes the fine-grained memory behavior of parallel workloads that demonstrates a reduction in EDP on a range of HPC benchmarks with minimal impact on execution time when compared to the standard OS/hardware-ma...
Nanda Velugoti, Joseph Manzano, Andrés Márquez et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.