Skip to content
Preprint

Density Estimation on Compact Manifolds under Intrinsic Spectral Block Variation

Aug 2026 · 0 citations
Mathematics

Abstract

We introduce an intrinsic spectral sparsity model for nonparametric density estimation on compact connected Riemannian manifolds. Instead of penalizing coefficients in an arbitrarily chosen Laplace--Beltrami eigenbasis, we group each complete eigenspace and measure the Hilbert norm of its spectral component. The resulting block-variation space is basis independent and isometry invariant. We establish its structural, atomic, and nonlinear approximation properties and clarify its relation to Sobolev, Besov, and coefficientwise spectral $\ell^1$ classes. We then construct a coordinate-free block-shrinkage estimator and prove a nonasymptotic signal-dependent $L^2$-oracle inequality that adapts to the unknown set of detectable eigenspaces. Under polynomial spectral growth, the risk theory separates the number of spectral blocks from their multiplicities and exhibits two regimes: one driven by a single high-dimensional eigenspace and the other by cumulative spectral complexity. Under matching spectral-growth and nondegeneracy assumptions, corresponding minimax lower bounds show that this multiplicity dependence is intrinsic, with sharp consequences for spheres and the rotation group $SO(3)$. Finally, we develop a positive, normalized, block-penalized exponential spectral sieve for log-densities and derive likelihood oracle inequalities together with expected Kullback--Leibler, Hellinger, and $L^2$ risk bounds. The resulting framework provides a geometry-respecting theory of sparse density estimation that remains invariant under changes of eigenbasis.

View source

Similar papers

Preprint Aug 2026

Spectral Dependence of Convex Regularization: Fundamental Limits under Right-Rotationally Invariant Designs

We study the fundamental limits of convex-regularized estimation in high-dimensional linear regression with right-rotationally invariant design matrices. We show that the asymptotic risks of convex-penalized least squares estimators are lower bounded by the risk of an approximate message passing algorithm known as Bayes VAMP, and we further characterize when the lower bound is attainable. As a technical ingredient in the proof of our lower bound theorem, we characterize the asymptotic performance of the $\ell_2$-perturbed convex estimator for every fixed perturbation strength $\lambda>0$. This closes a gap in the literature, in which $\lambda$ was required to be sufficiently large or restrictions were imposed on the class of convex estimators. The benefit of our approach is that we can conduct a direct analysis of the spectrum's impact on the lower bound in the high-dimensional limit. In particular, we can show that the lower bound is monotone in an ordering on the limiting spectral distributions of the design matrix. These results isolate how the full singular-value distribution of the design---not merely the aspect ratio or average measurement strength---governs the limitations of convex regularization.

B. Tan, Audrey Yang, Cynthia Rush · 0 citations
Open access Aug 2026

Asymptotic Theory for Kernel Density Estimation Under Dependent Length-Biased Sampling

We establish an asymptotic theory for the Jones inverse-weighted kernel density estimator when length-biased observations form a strictly stationary short-range dependent sequence. The statistical difficulty is intrinsically composite: reciprocal weighting is singular at the origin, the normalizing mean is estimated from the same dependent sample, kernel localization shrinks with the bandwidth, and the centered summands form a row-wise stationary triangular array whose envelope diverges at rate hn−1. Under a non-negative compactly supported Lipschitz kernel, an inverse-moment condition, geometric α-mixing, local regularity of the target density, and uniform local bounds on lagged bivariate densities, we prove strong uniform consistency on compact subsets of (0,∞) and, separately, the uniform stochastic bound OP{hn2+(logn/(nhn))1/2}. A covariance-localization argument shows that the scaled serial-covariance contribution is O{hnlog(1/hn)}=o(1), so the first-order pointwise variance coincides with that of the corresponding independent length-biased estimator. Pointwise and finite-dimensional Gaussian limits are obtained by an explicit big-block/small-block argument with off-diagonal covariance control. The ratio normalization is treated directly: its variance contribution, its product with the localized fluctuation, and its cross-covariance with that fluctuation are all negligible at the nhn scale. We further derive first-order AMSE and AMISE criteria, their oracle bandwidths, and feasible pointwise studentization under undersmoothing. The numerical study separates oracle from data-driven bandwidth selection, evaluates full-ratio HAC and moving-block corrections, examines a Frank-copula Markov robustness design, and benchmarks the Jones estimator against an alternative length-biased estimator. The simulations support the first-order theory while demonstrating that persistent short-range dependence can remain consequential for finite-sample uncertainty.

Salim Bouzebda, S. Didi · 0 citations
Preprint Aug 2026

High-Dimensional Spectral Limits for Gaussian KL-Unbalanced Optimal Transport

We study high-dimensional random-matrix limits of Gaussian Kullback--Leibler unbalanced optimal transport (KL-UOT). Under equal marginal penalties, the covariance action admits an exact log-determinant representation in terms of a nonlinear ridge product, together with a positive-semidefinite extension that remains finite at arbitrary aspect ratios. For independent real Wishart samples, strong asymptotic freeness gives the limiting free multiplicative convolution and almost-sure Hausdorff convergence of the ridge-product spectrum; independent Haar orientations yield the corresponding first-order limit for deformed populations. In the symmetric nonsingular identity-Wishart model, we derive an explicit $\eta$-transform and a low-degree algebraic equation that select the physical branch and determine the support interval, square-root edges, and extreme-eigenvalue limits. We further obtain all-aspect one-sample Marchenko--Pastur limits under finite fourth moments, real-Gaussian Bai--Silverstein fluctuations for $c<1$, and a joint random-matrix/penalty limit showing that sample-covariance noise produces the critical scale $\tau_p\asymp p$.

Jia-Ping Yang, Yun-Xin Zhang · 0 citations
Preprint Sep 2026

On the sample complexity of the active subspace method

Active subspaces identify low-dimensional linear structure in high-dimensional parameter-to-output maps by estimating the dominant eigenspace of a gradient covariance operator. In practice this covariance is replaced by a Monte Carlo estimator built from a limited number of gradient evaluations. Classical analyses based on controlling the covariance error in operator norm lead to sample-complexity estimates that can be substantially more pessimistic than the sampling rules commonly used in computations. This paper studies the empirical active subspace method directly in the projection-error metric relevant for ridge approximation. We derive non-asymptotic quasi-optimality bounds governed by a regularized inverse Christoffel function associated with the gradient field. Under a bounded-gradient assumption, the resulting estimates already improve the sample-complexity estimates obtained from operator-norm covariance bounds. We then show that additional smoothness of the gradient map, expressed through membership in a reproducing kernel Hilbert space, yields sharper coherence estimates and motivates tractable importance sampling from kernel diagonal measures. Furthermore, the same smoothness assumption yields a priori decay bounds for the population active subspace tail energy, which can be combined with our finite-sample estimate to prescribe rank, regularization scale, and sample size, allowing to fully characterize the a priori sample complexity. The abstract assumptions are verified for lognormal Gaussian and affine uniform parametric elliptic PDEs using weighted summability of Hermite and Legendre series expansions.

Fabio Nobile, Matteo Raviola, R. Tempone · 0 citations
Preprint Aug 2026

Posterior contraction rates in Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families

We study posterior contraction in positive-order Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families. We embed the natural parameter in a Hilbert scale and model it via a standard Gaussian series prior expanded in the eigenbasis generating the scale. Under a two-sided link condition on the Fisher information and suitable local regularity assumptions, we show that smoothness-matching priors achieve minimax-optimal posterior contraction rates in any Hilbert scale norm up to the regularity of the ground truth. Our analysis builds on the novel approach to posterior contraction based on the Wasserstein distance recently introduced by Dolera et al. (2024). It combines refined Laplace-type estimates for infinite-dimensional integrals associated to the posterior kernels with a mixed-geometry estimate controlling their stability under fluctuations in the data, itself resting on a tailored Poincar\'e inequality for posterior distributions conditioned on neighbourhoods of the truth. We apply the general theory to density estimation with a logistic parametrisation, Poisson intensity estimation with an exponential link, and the Gaussian white-noise model, yielding minimax contraction rates in Sobolev norms across all three settings. In particular, these yield optimal recovery of density score functions and derivatives of Poisson intensities.

Emanuele Dolera, Stefano Favaro, M. Giordano · 0 citations
Preprint Sep 2026

Covariance and Principal Component Analysis on Riemannian Manifolds and Graphs

We develop notions of covariance and principal component analysis (PCA) for probability measures on Riemannian manifolds of bounded geometry and a discrete counterpart for distributions on the vertex sets of weighted simple graphs. Rather than anchoring variation at the Fr\'{e}chet mean, whose uniqueness and usefulness can fail in this context, we consider variation about every point. This leads to a representation of each point on the manifold by a vector field derived from the heat kernel and thus to a map of the underlying manifold into a Hilbert space of vector fields, in which covariance and PCA carry over from the Euclidean setting. A distinguishing feature of this formulation is that one typically obtains infinitely many principal components, capable of capturing highly nonlinear geometric features and patterns of variation. We prove an embedding theorem for the mapping into the space of vector fields and develop two computational reductions of the theory: RiePCA, a finite-dimensional reduction based on vector fields supported on finitely many points, and GraphPCA, a discrete formulation for measures on weighted graphs in which vector fields assign orientations and magnitudes to edges. Numerical examples and experiments illustrate the behavior of the proposed methods on both synthetic and real data.

L. Hartmann, Wen-Wen Li, W. Mio · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.