Convergence Guarantees of Gradient Descent for Neural Networks via Generalized Lipschitz Smoothness
We establish the first convergence guarantees of gradient descent for general feedforward neural networks of any width or depth, with any initialization or dataset. We only assume that the activation functions are linearly bounded, Lipschitz continuous, and Lipschitz smooth---properties that hold for linear, tanh, soft...