Skip to content

Amortized Bandwidth Learning for Kernel Density Estimation under Logarithmic Score

Aug 2026 · 0 citations · 20 references
Computer Science Mathematics

TL;DR

An amortized framework is proposed that consistently and substantially outperforms Silverman's rule, the Sheather--Jones selector, and least-squares cross-validation, with especially large gains in small and heterogeneous samples.

Abstract

Kernel density estimation converts finite samples into probability densities, but its performance depends critically on bandwidth selection. Classical selectors prescribe the sample-to-bandwidth rule analytically or asymptotically, or solve a new optimization for each sample. An amortized framework is proposed that instead learns this mapping across a distribution of density-estimation tasks by optimizing the logarithmic score. A truncated-and-renormalized bounded-support formulation enables stable learning across heterogeneous tasks, while affine standardization allows a selector trained on a single reference interval to transfer across bounded intervals. Experiments under Gaussian sampling, a multi-family benchmark, and randomized Gaussian-mixture training show that the amortized selector consistently and substantially outperforms Silverman's rule, the Sheather--Jones selector, and least-squares cross-validation, with especially large gains in small and heterogeneous samples. Finite Gaussian mixtures provide a generic training mechanism supported by their $L^1$ approximation property. Selectors trained in this way generalize strongly across different density structures, allowing the same trained selector to be applied directly to finite samples from unknown densities without specifying or fitting a distributional family. This combination of broad applicability and strong empirical performance makes the framework attractive for a wide range of applications in which finite samples or ensembles must be converted into continuous probability densities.

View source

Similar papers

#machine learning Preprint Sep 2026

Learning under Target Shift: Optimal Density Ratio Estimation and Importance-Weighted Regression

We study density ratio estimation and importance-weighted regression under target shift with continuous outputs. Under target shift, the conditional distribution of the inputs given the outputs remains invariant across the training and test distributions, while the output marginal distribution may change. Although this problem has been extensively studied for discrete outputs, the continuous setting is substantially less understood: the importance weights are determined by an unknown density ratio function, for which existing estimation methods lack explicit finite-sample convergence rates. We propose a spectral regularization method in a reproducing kernel Hilbert space (RKHS) for estimating the continuous density ratio from labeled training samples and unlabeled test inputs. Under a source condition with regularity parameter $\iota>0$, we establish high-probability finite-sample guarantees and show that the estimator achieves the capacity-independent minimax-optimal RKHS-norm rate $O(n_\eta^{-\iota/(2\iota+2)})$. We then incorporate the estimated density ratio into importance-weighted regression and characterize the propagation of density-ratio estimation error to the final predictor. When sufficiently many samples are available for density ratio estimation, the resulting regression estimator attains the minimax-optimal rates of standard kernel regression. These results establish a finite-sample theory for continuous density ratio estimation and importance-weighted learning under target shift.

Ren-Wei Liu, Zheng-Chu Guo · 0 citations
Open access Dec 2025

Learning Lévy density via adaptive RKHS regression with bi-level optimization

Numerical experiments demonstrate that the bilevel RKHS method provides a more stable and competitive alternative to classical L-curve and generalized cross-validation strategies and that the adaptive RKHS norm is more accurate and robust than Lρ2- and ℓ2-norms for regularization.

Luxuan Yang, Fei Lu, Ting Gao et al. · 0 citations
Open access Aug 2026

Asymptotic Theory for Kernel Density Estimation Under Dependent Length-Biased Sampling

We establish an asymptotic theory for the Jones inverse-weighted kernel density estimator when length-biased observations form a strictly stationary short-range dependent sequence. The statistical difficulty is intrinsically composite: reciprocal weighting is singular at the origin, the normalizing mean is estimated from the same dependent sample, kernel localization shrinks with the bandwidth, and the centered summands form a row-wise stationary triangular array whose envelope diverges at rate hn−1. Under a non-negative compactly supported Lipschitz kernel, an inverse-moment condition, geometric α-mixing, local regularity of the target density, and uniform local bounds on lagged bivariate densities, we prove strong uniform consistency on compact subsets of (0,∞) and, separately, the uniform stochastic bound OP{hn2+(logn/(nhn))1/2}. A covariance-localization argument shows that the scaled serial-covariance contribution is O{hnlog(1/hn)}=o(1), so the first-order pointwise variance coincides with that of the corresponding independent length-biased estimator. Pointwise and finite-dimensional Gaussian limits are obtained by an explicit big-block/small-block argument with off-diagonal covariance control. The ratio normalization is treated directly: its variance contribution, its product with the localized fluctuation, and its cross-covariance with that fluctuation are all negligible at the nhn scale. We further derive first-order AMSE and AMISE criteria, their oracle bandwidths, and feasible pointwise studentization under undersmoothing. The numerical study separates oracle from data-driven bandwidth selection, evaluates full-ratio HAC and moving-block corrections, examines a Frank-copula Markov robustness design, and benchmarks the Jones estimator against an alternative length-biased estimator. The simulations support the first-order theory while demonstrating that persistent short-range dependence can remain consequential for finite-sample uncertainty.

Salim Bouzebda, S. Didi · 0 citations
Preprint Aug 2026

Fast high-dimensional mean testing via logistic regression

We propose computationally efficient tests for equality of mean vectors of two or more high-dimensional populations. Central to our approach is an equivalence between equality of means and a zero population logistic regression parameter. We establish this equivalence for independently distributed observations without imposing common distributional assumptions across populations. Our procedure uses logistic Lasso to screen informative variables and an unpenalized logistic refit for inference in the reduced dimension, yielding asymptotically correct size and consistency. For a specified two-sample Gaussian submodel and sparse discriminative class, the test also attains the minimax separation rate. The framework extends to multiple populations through multi-class logistic regression. Simulations demonstrate accurate size control, strong power, and favorable computational scaling compared with existing tests under unbalanced designs and variance heterogeneity. Applications to gene-expression data with more than twenty-two thousand variables illustrate the practical scalability of the proposed procedures.

Sayan Das, Debraj Das, S. Dutta · 0 citations
Preprint Aug 2026

Nonparametric Goodness-of-fit Testing under Covariate Shift

This paper develops procedures for nonparametric goodness-of-fit testing under covariate shift, where labelled data are drawn from a source population but goodness-of-fit is evaluated for a target population. The distribution mismatch is quantified by either a bounded moment condition or a sub-exponential tail condition on the target-to-source density ratio. Our method combines truncated importance-weighting kernel ridge regression with a multiplier bootstrap to construct confidence sets for the regression function. The truncation stabilizes the importance- weighting kernel ridge regression as well as the bootstrap calibration, making our approach applicable even when the density ratio has heavy tails. We prove nonasymptotic validity and sharpness of the resulting confidence sets under suitable operator compatibility conditions, and establish explicit error rates for coverage probability under specific conditions on the target- to-source density ratio and on the spectral decay of the kernel integral operator. Numerical experiments corroborate our theoretical findings.

Zhengyou Hou, Dong Xia · 0 citations
Preprint Sep 2026

Covariate-localized False Discovery Rates

We introduce a flexible model for covariate-dependent multiple testing which can be encoded using a nonparametric Gaussian mixture model. Weight-localized predictive recursion (PRx), a new development in the methodology of Newton's predictive recursion algorithm, is then leveraged to estimate the components of this mixture model, allowing for recovery of the covariate-localized false discovery rate $\text{Pr}(H_i = 0|z_i,x_i)$ using a single, unified algorithm. This quantity represents the most direct extension of Efron's local false discovery rate to the covariate-dependent setting, and admits provable Bayesian FDR control properties under simple rejection rules. We introduce several procedures for estimating and thresholding the local false discovery rate, and show using various simulations and a real-data example that our procedures lead to increased power, tighter Bayesian FDR control, and more interpretable rejections. We furthermore show that this holds for fixed and randomized hypothesis labels, indicating that our proposed methods perform well under both frequentist and Bayesian interpretations of multiple testing.

Jonathan Lin, S. Tokdar · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.