Low-bit and Sparsified Gradient Communication for Accelerating Distributed Deep Learning with Convergence Guarantees
Communication poses a dominant bottleneck in distributed data parallel training with synchronous stochastic gradient descent, whereas the traditional AllReduce collective used for gradient synchronization limits the efficient utilization of communication compression strategies. In this paper, we propose a low-bit and s...