Skip to content

Author

Shutai Yang

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Optimal Parameter-Free Gradient Minimization in $\ell_p$ Geometry

We study the first-order oracle complexity of finding a queried point with small gradient in $\ell_p$ geometry, with particular attention to the information needed to adapt the unknown smoothness and distance scales. In the strict counted local value--gradient model, no finite complexity bound can depend only on $LR/\eps$ without a nondegenerate local scale observation: a one-dimensional construction keeps $LR/\eps=4$ while defeating every prescribed finite query budget. We resolve Diakonikolas's general-$\ell_p$ parameter-free extension question for every fixed $1<p<\infty$. Under a nondegenerate secant initialization, the method knows neither the smoothness constant $L$, the initial solution distance $R$, nor $f^*$, and returns a queried point $\widehat x$ with $\|\nabla f(\widehat x)\|_q\le\eps$. For fixed finite $p>2$, we first establish the dimension-free deterministic known-parameter upper exponent $p/(p+2)$ in $K=LR/\eps$, matching the published lower polynomial exponent under its horizon and dimension qualifications. The finite local routine fits the same observable scale--radius procedure, so this exponent is preserved without knowing $L$ or $R$. Writing $\Kbar=\max\{1,LR/\eps\}$, the post-initialization pair-oracle complexity is $O_p(\Kbar^{1/2})$ for $1<p<2$, $O(\Kbar^{1/2})$ for $p=2$, and $O_p(\Kbar^{p/(p+2)})$ for $p>2$, together with the additive calibration cost $O_p(\log(e+L/M_0))$ in every regime.

Shutai Yang, Yu-Ning Yang · 0 citations
Preprint Aug 2026

The Sharp Worst-Case Asymptotic Rate of the Barzilai--Borwein Method in $\mathbb R^d$ and Hilbert Spaces

We establish sharp asymptotic rates for the two Barzilai--Borwein (BB) rules on uniformly positive quadratics and local nonlinear problems. In finite dimensions, for either fixed rule and an arbitrary positive first step, the gradient root factor is bounded by $(b_0-a_0)/(b_0+a_0)$, where $[a_0,b_0]$ is the initially active spectral interval. Hence the worst trajectory factor is $c_H=(\kappa(H)-1)/(\kappa(H)+1)$. When $H$ has at least two distinct eigenvalues, matched initialization and a balanced endpoint trajectory attain this value. Under matched initialization, the same constant is the optimal uniform-envelope threshold. For bounded, self-adjoint, uniformly positive operators on Hilbert space, scalar spectral measures yield the corresponding active-support bound and optimal matched uniform-envelope threshold, including continuous endpoint spectrum. Finally, if the gradient is strictly Fr\'echet differentiable at a stationary point and its derivative is self-adjoint and uniformly positive, every $\gamma\in(c_*,1)$, where $c_*=(\kappa(A_*)-1)/(\kappa(A_*)+1)$, is a uniform local envelope rate for either pure BB rule. Every well-defined trajectory converging to the stationary point has error and gradient root factors at most $c_*$ and objective-gap root factor at most $c_*^2$. Over the class of objectives with prescribed distinct derivative endpoints $m_*

Shutai Yang, Ya-xiang Yuan · 0 citations