Chapel for Parallel AND Distributed GPU Computing: A Case Study with Jaccard Similarity
Overall, when evaluating both Chapel and MPI+X implementations of partitioned Jaccard similarity on up to 16 A100-80GB GPUs across four nodes, Chapel achieves comparable performance to MPI+X while also delivering significantly better programmer productivity and agility.
Paul Sathre, Wu-Chun Feng
· Proceedings of the Internati... · 0 citations