Revisiting AdaGrad in Stochastic Convex Optimization: Last Iterates, High Probability, and Lower Bounds
It is proved that bounded variance alone is too weak: even with bounded iterates, it cannot yield high-probability average-iterate rates, and without bounded iterates it may not even guarantee convergence in expectation.