Optimizing point-to-point and collective communication in HPC systems substantially impacts the performance of large-scale applications. Middleware libraries such as Message Passing Interface (MPI) rely on lower-level communication libraries, including UCX and OFI Libfabric, to enable these optimizations. However, because optimal configurations depend on the characteristics of each HPC system, users must understand system-specific performance behavior and manually tune communication parameters. In this work, we present HAT-MPI, an ML-based framework that identifies the UCX Eager-Rendezvous protocol threshold and the optimal collective algorithm on InfiniBand HPC clusters. Using hardware characteristics and performance data collected across clusters at different scales, HAT-MPI predicts the UCX protocol threshold and the optimal Alltoall/Allreduce algorithm for previously unseen InfiniBand clusters. This paper also introduces two complementary techniques that improve the conventional data collection and evaluation workflow, namely LLM-assisted HPC hardware detection and Top-5% candidate algorithm evaluation. Experimental results show that the predicted UCX protocol threshold achieves an average of 31% performance improvement in regions where a gap exists between the predicted and default thresholds, while the algorithm selection attains 84.9% accuracy with Top-5% evaluation, demonstrating that HAT-MPI delivers highly accurate predictions.
Sungjae Lee, Sam Tilford, D. Panda· IEEE International Symposium...· 0 citations
Preliminary results from a geo-distributed LLM training prototype that treats networking constraints as first-order design concerns are presented, motivating adaptive networking support for synchronization, compression, placement, telemetry, and recovery in geo-distributed LLM training.
Ziyue Luo, Jiaxuan Cai, Cedric Le Denmat et al.· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.