Skip to content
Preprint

Spectra of high-dimensional Spearman correlation matrices under scale-mixture dependence

Jul 2026 · 0 citations · 27 references
Mathematics

Abstract

We study the asymptotic spectral properties of high-dimensional Spearman correlation matrices for scale-mixture data. We consider observations of the form $x_t=\sigma_t \xi_t \in \mathbb{R}^N,$ where the coordinates of $\xi_t$ are i.i.d.\ and the scalar mixture variable $\sigma_t$ is shared by all coordinates. Under natural symmetry assumptions, the coordinates of $x_t$ are pairwise uncorrelated in both the Pearson and Spearman sense. Nevertheless, they are not independent when the mixture variable is non-degenerate. We show that this higher-order dependence survives the rank transformation and leaves a nontrivial spectral signature. In the proportional regime $N/T\to q\in(0,\infty),$ the empirical spectral distribution of the Spearman correlation matrix converges almost surely to a generalized Mar\v{c}enko--Pastur law governed by the limiting distribution of an effective rank variance. We also formulate a broader latent-variable extension, which covers, in particular, some scale-mixture models with correlated directional components. We discuss solvable examples and numerical approximations, motivated in part by heavy-tailed data in robust multivariate statistics, econometrics, and finance.

View source

Similar papers

Preprint Aug 2026

Non-Gaussian fluctuations for traces of squared sample correlation matrices in high dimensions

We provide limit theory for the trace of the squared sample correlation matrix $\mathbf R$, constructed from $n$ observations of a $p$-dimensional random vector with iid components. If the entries have finite fourth moment and $p$ and $n$ grow proportionally, it is known that $\operatorname{tr}({\mathbf R}^2)$ satisfies a central limit theorem (CLT) and the centering and scaling sequences are universal in the sense that they do not depend on the entry distribution. Under symmetry and regular variation assumption with index $\alpha$ and any growth rate of the dimension, we prove that the universal CLT remains valid for $\alpha>3$. For $\alpha<3$, we identify a critical dimension growth at which the fluctuations of $\operatorname{tr}({\mathbf R}^2)$ become non-Gaussian. Moreover, if the dimension $p$ grows faster and $\alpha\le 3$ we establish a non-universal CLT with norming sequences depending on the value of $\alpha$. Our findings are illustrated in a simulation study.

J. Heiny, Xuechun Hu, Felix J. Seo · 0 citations
Preprint Sep 2026

Phase transition for the smallest eigenvalue of high-dimensional sample correlation matrices

We study the smallest nonzero eigenvalue of the sample correlation matrix $\mathbf{R}_n$ formed from a $p_n \times n$ data matrix with i.i.d. real entries $\xi$ of mean zero and unit variance, in the high-dimensional regime $p_n / n \to \phi \in (0, \infty) \setminus \{1\}$. In the tall regime $\phi>1$, we prove almost-sure convergence of $\lambda_n (\mathbf{R}_n)$ to the lower Mar\v{c}enko--Pastur edge $\lambda_- = (1 - \sqrt{\phi})^2$ without additional moment assumptions. In the wide regime $\phi<1$, we establish a phase transition at the third-order tail scale. If $t^3 \mathbb{P}\{\lvert \xi \rvert>t\} \to 0$ as $t \to \infty$, the smallest eigenvalue $\lambda_{p_n} (\mathbf{R}_n)$ converges in probability to $\lambda_-$. While if $t^3 \mathbb{P}\{\lvert \xi \rvert>t\} \to \infty$, then $\lambda_{p_n} (\mathbf{R}_n)$ converges in probability to zero. At the critical scale $t^3 \mathbb{P}\{\lvert \xi \rvert>t\} \to \kappa \in (0, \infty)$, the point process of eigenvalues in the lower gap $(0, \lambda_-)$ converges in distribution to a Poisson random measure with explicit intensity. In this critical regime, we also identify the nondegenerate limiting distribution of $\lambda_{p_n} (\mathbf{R}_n)$, which has a continuous density on $(0, \lambda_-)$ and a positive atom at $\lambda_-$.

Ze-Qin Lin, Guamgming Pan, Hao-Zhu Zhao et al. · 0 citations
Preprint Aug 2026

Correlation Matrices in High Dimensions: The Elliptope as a Sample-Correlation Ensemble

The set of $n\times n$ correlation matrices, known as the elliptope, has volume decaying at the super-exponential rate $\exp\{-\tfrac14 n^2\log n\}$. We characterize where this vanishing volume concentrates. A uniform draw is entrywise close to the identity yet globally far from it and nearly singular: its maximum absolute correlation is of order $\sqrt{\log n/n}$, its Frobenius distance is asymptotic to $\sqrt n$, its empirical spectral distribution converges to the Marchenko-Pastur law with ratio one, and its smallest eigenvalue has the exact $\operatorname{Beta}(1,d)$ distribution, where $d=n(n-1)/2$, and is therefore of order $n^{-2}$. More generally, distinct off-diagonal entries are exactly pairwise independent under every $\operatorname{LKJ}(\eta)$ law. For the uniform law, this yields a Chen-Stein proof of the extreme-correlation point-process limit and an $O(n^{-1})$ total-variation bound for finite-dimensional exceedance counts relative to Poisson laws with their exact finite-$n$ means. We also identify two distinct scales: $\eta_n\asymp n$ alters the limiting spectrum, whereas $\eta_n\asymp n^2$ is needed to keep the Frobenius distance bounded. Finally, for a bounded, centered i.i.d. off-diagonal specification, projection to the nearest correlation matrix incurs a squared repair cost asymptotically at least one-half of the squared Frobenius norm of its off-diagonal part.

P. Hansen · 1 citation
Preprint Jul 2026

The Phase Transition in Online PCA Depends on $n/d\log(d)$, not $n/d$

High dimensional statistical theory has established the importance of constant aspect ratio, when the number of dimensions ($d$) and samples ($n$) satisfy $n,d\to\infty$ with $n/d\to \gamma\in(0,\infty)$, in understanding the limits of canonical estimation problems. In particular, for estimating the top eigenvector of a $d\times d$ population covariance matrix from $n$ iid samples, the BBP phase transition gives a precise threshold -- a simple functional of the aspect ratio -- such that the top sample principal component attains nonzero asymptotic correlation with the truth only when the leading population eigenvalue exceeds it. In this paper, we show that for online / streaming algorithms the story is very different, and constant aspect ratio is insufficient for nonzero overlap. We study Oja's algorithm, the most popular method for online PCA. Let $\Sigma=\theta^2 v_0v_0^\top+I\in\mathbb{R}^{d\times d}$, and run Oja's algorithm with step size $\delta/d$ on $n$ iid samples $X_k\sim\mathcal{N}(0,\Sigma)$, with output $\hat v_n$. Then, as $n,d\to\infty$ with $n/d\log d\to\gamma\in(0,\infty)$, we establish a phase transition: $|\langle\hat v_n,v_0\rangle|\to 0$ when $\gamma<\gamma_*$, and $\to\rho_*$ when $\gamma>\gamma_*$. Here $\rho_*=\rho_*(\theta,\delta)=\sqrt{(\theta^2-\delta/2)_+/\theta^2(1+\delta/2)}$ and $\gamma_*=\gamma_*(\theta,\delta)=1/2\delta(\theta^2-\delta/2)_+$. Further, at criticality, when $n=[\gamma_*d\log d+\eta d]$ and $d\to\infty$, $\eta\in\mathbb{R}$, the correlation is random: $|\langle\hat v_n,v_0\rangle|\stackrel{w}{\to}\rho_*|G|\exp(\eta/2\gamma_*)/\sqrt{\rho_*^4+G^2\exp(\eta/\gamma_*)}$ where $G\sim\mathcal{N}(0,1)$. This is in stark contrast to ordinary high dimensional PCA, where nonzero overlap is possible at constant $n/d$ and improves as $n/d$ increases.

Apratim Dey · 0 citations
Preprint Sep 2026

Geometric Fluctuations of the $\sin\Theta$ Distance in High-Dimensional Principal Subspace Estimation

We investigate the geometric fluctuations of principal subspaces for high-dimensional covariance matrices through the squared Frobenius $\sin\Theta$ distance between the sample and population eigenspaces associated with the $r_p$ largest eigenvalues. An explicit first-order expansion and a central limit theorem are established for this subspace distance. The theory allows the subspace dimension to diverge subject to $r_p=o(n)$, where $n$ is the sample size. It also permits a diverging spectral norm of the population covariance matrix, population spikes of different orders, and repeated or closely spaced spikes. This sharp characterisation captures features of the subspace estimation error that are not reflected in existing perturbation bounds. As applications, we derive an explicit asymptotic expansion for the expected PCA excess risk and a refined error bound for distributed PCA. In both cases, existing upper bounds can increase with the spiked-block condition number when some leading spikes become stronger, whereas our results show that the corresponding estimation errors need not increase and may instead decrease. Numerical experiments reproduce this contrasting behaviour and demonstrate the finite-sample accuracy of our theoretical findings.

Yan-Ling Hu, Xiao Han, Qing Yang · 0 citations
Preprint Sep 2026

Sharp spectral norm concentration of sparse random tensors

We prove a sharp concentration inequality for the spectral norm of sparse random tensors with independent Bernoulli entries. Let $T$ be an order-$k$ tensor of dimension $n\times\cdots\times n$ with independent Bernoulli$(p)$ entries, where $k$ is fixed. For any $c,r>0$, we show that $\|T-\mathbb E T\|\le C_{k,r,c}\sqrt{np}$ with probability at least $1-n^{-r}$ whenever $np\ge c\log n$. We extend this bound to inhomogeneous Bernoulli sampling with deterministic entrywise weights. This removes the logarithmic factor in the work of Zhou and Zhu (2021). The proof follows the Kahn--Szemer\'edi light--heavy decomposition with a refined estimate on the heavy tuple part. We also obtain a log-free second eigenvalue bound for the random hypergraph model of Friedman and Wigderson (1995).

Zhixin Zhou, Yizhe Zhu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.