BAG: Faster Matrix Multiplication on a Single GPU
BAG (Basis Alternative Matrix Multiplication on GPUs), a GPU-oriented implementation of ABMM for the NVIDIA Ampere architecture is presented, designed to shrink workspace and eliminate redundant global-memory traffic, and introduce a cost-model-based recursion policy together with a Roofline-guided blocking strategy to...