Skip to content
Preprint

Upper Confidence Bounds for the Prediction Error of Kernel Ridge Regression via Gaussian Refitting

Jul 2026 · 2 citations · 45 references
Mathematics

TL;DR

This work proposes a Gaussian refit for kernel ridge regression by Anderson's inequality, which requires no moment assumptions and is calibrated at any confidence level via order statistics, and extends empirically to nonlinear constrained estimators and real spatial data.

Abstract

Assessing a single model fit requires a computable upper confidence bound for the gap between the fit and the unknown truth, as mean estimates ignore realization variance. Standard cross-validation margins are bottlenecked at order $n^{-1/2}$ by noise fluctuations, even when the true error shrinks faster. While wild refitting cancels this noise level, existing Rademacher sign methods degenerate for kernel ridge regression and rely on unobservable quantities. We propose a Gaussian refit for kernel ridge regression. By Anderson's inequality, the fit movement is monotone in the noise sizes, yielding a computable tail bound. Assuming only symmetric noise, the bound requires no moment assumptions and is calibrated at any confidence level via order statistics. Theoretically, using a worst-case envelope, the bound contracts at the minimax rate $O_P(n^{-2s/(2s+1)})$, correctly matching the prediction error. Empirically, using a practical data-driven envelope, the bound maintains full coverage within twice the true $95\%$ error quantile. By contrast, cross-validation exceeds this quantile by factors up to $51$, and by hundreds under infinite-variance noise. The procedure extends empirically to nonlinear constrained estimators and real spatial data.

View source

Similar papers

Preprint Jul 2026

Exact Generalization Error Curves of Kernel Ridge Regression for Functional Moment Estimation

Kernel ridge regression is a standard method for functional data analysis, but its exact behavior is less understood. We study tensor-product kernel ridge regression for estimating the $r$-th moment function of a random function based on noisy discrete observations. The formulation includes mean estimation, covariance...

Yi Ding, Yicheng Li · 0 citations
Preprint Aug 2026

High-dimensional ridgeless least squares interpolation under spiked covariance structures

This paper investigates the asymptotic behavior of the out-of-sample prediction risk of the high-dimensional ridgeless least-squares estimator when the feature dimension $p$ and the sample size $n$ grow proportionally. We consider a generalized spiked population covariance model with multiple latent factors, where the...

Zhi-Jun Liu, Dandan Jiang · 0 citations
Preprint Sep 2026

Covariate-localized False Discovery Rates

We introduce a flexible model for covariate-dependent multiple testing which can be encoded using a nonparametric Gaussian mixture model. Weight-localized predictive recursion (PRx), a new development in the methodology of Newton's predictive recursion algorithm, is then leveraged to estimate the components of this mix...

Jonathan Lin, S. Tokdar · 0 citations
Preprint Jul 2026

The V-fold jackknife for semiparametric inference: variance estimation, confidence intervals, and simultaneous confidence bands

For decades, the bootstrap has been a default tool for statistical inference because of its broad applicability and minimal analytic requirements. Although its validity is well understood for smooth parametric estimators, its theoretical properties for many modern semiparametric and machine-learning estimators remain l...

Yi Li, Ashkan Ertefaie, M. J. van der Laan · 0 citations
Preprint Jul 2026

Amortized Inference for Sampling Distributions Where the Bootstrap Fails

A neural network is trained on simulated datasets drawn from a prior over a distribution family, using single independent draws of the root T_n - T(F) scored by the pinball loss, a proper scoring rule whose population minimizer is the posterior-predictive law of the root.

Akash Deep · 0 citations
Preprint Sep 2026

Generalized Ridge Refitting for the Lasso and Prediction Improvement Bounds

We study a class of Lasso based estimators obtained by applying a quadratic correction on the Lasso equicorrelation set. The penalty matrix determines both the magnitude and geometry of the correction and contains, among other cases, the isotropic Lasso--Ridge correction, least squares refitting, Gram proportional inte...

Guo Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.