Exhaustive moment fitting in this constant-dimensional space produces a proper mixture and, together with the dimension-free moment characterization of Gaussian mixtures, achieves the optimal Hellinger rate in polynomial arithmetic time for every fixed $k$.
Abstract
We consider a mixture of at most $k$ unit-covariance Gaussians in $\mathbb{R}^d$ whose means belong to a fixed-radius ball, with no separation or minimum-weight condition. Doss, Wu, Yang and Zhou (2023) proved that the minimax Hellinger risk is of order $\sqrt{d/n}\wedge 1$ and constructed a proper polynomial-time estimator with the slower general bound $(d/n)^{1/4}$; obtaining the sharp rate in polynomial time for fixed $k\geq 3$ was left open. We resolve this question. The key device is a moment-fiber range finder. A second-moment subspace controls the energy missed by projection. We then estimate finitely many one-free-index Hermite contractions. These vector-valued contractions recover every tensor component containing exactly one missed direction at the sharp $\sqrt{d/n}$ scale. Every remaining term contains at least two missed factors and is therefore controlled by the residual second-moment energy. The resulting subspace has dimension depending only on $k$. Exhaustive moment fitting in this constant-dimensional space produces a proper mixture and, together with the dimension-free moment characterization of Gaussian mixtures, achieves the optimal Hellinger rate in polynomial arithmetic time for every fixed $k$.
Let $X$ be a centered, variance-one random variable with finite fourth moment, and form the principal degree-$d$ tensor feature vector of all square-free monomials in $n$ independent copies of $X$. For $m$ independent samples we determine the global spectral law of the sample covariance throughout the critical scale $d^2/n\to\lambda\in[0,\infty)$, with aspect ratio $p/m\to c$. For a fixed base distribution with finite fourth moment and $P(|X|=1)<1$, prior work gives ordinary Marchenko--Pastur convergence if and only if $d=o(\sqrt n)$. We identify the finite critical boundary: when $d^2/n\to\lambda\in(0,\infty)$, the tensor radius converges in quadratic Wasserstein distance to a lognormal law determined by the fourth moment, while all remaining bounded quadratic fluctuations vanish. A leave-one-out resolvent argument then yields almost-sure convergence of the empirical spectral distribution to a free compound-Poisson law driven by this endogenous lognormal jump. The limit reduces to Marchenko--Pastur when the fourth-moment excess or the overlap intensity vanishes. In the unit-modulus case, our estimates recover the sharp range $\min(d,n-d)=o(n)$ for uniform quadratic-form concentration and imply Marchenko--Pastur convergence throughout that range, with an explicit variance bound.
We study the problem of recovering the correspondence between a collection of $n$ points in $\mathbb{R}^d$ and a noisy, permuted version of those points. In the high-dimensional regime $d=\omega(\log n)$, under a Gaussian model with noise variance $\sigma^2=d/(b\log n)$, prior work identifies $b=2$ as the threshold for almost exact recovery. We prove that this threshold is all-or-nothing: for every fixed $b<2$, no estimator recovers a positive fraction of the matching, and even estimating the matched point cloud in Euclidean distance is asymptotically no better than ignoring the correspondence. On the other hand, we consider a multi-view generalization of the problem where $K$ noisy, independently permuted copies of the same latent point cloud are observed. Here we show that a simple polynomial-time procedure recovers all relative matchings up to $o(n)$ errors whenever $b>K/(K-1)$. Thus multiple views can break the impossibility barrier $b=2$ for the original matching problem: in particular, for $3/2<b<2$, the two-view model has no nontrivial recovery, but a third view makes all latent correspondences efficiently recoverable.
Timothy L. H. Wee, Kaylee Yingxi Yang, Zhou Fan et al.· 0 citations
We give explicit complex polynomials $P,Q$ in three independent standard real Gaussian variables such that \[ {\mathbb E}(P^m)=0,\qquad {\mathbb E}(QP^m)=m!\neq0 \] for every $m\geq1$. In natural complex linear coordinates, $P$ has five terms and total degree $4$. Hence the Gaussian Moments Conjecture is false in every dimension $n\geq3$. We also give a six-term cubic example in four variables, which was found first and already proves failure for every $n\geq4$. Both examples follow from the same coefficient identity. The search was prompted by Levent Alp\"oge's public announcement of an explicit three-dimensional counterexample to the Jacobian Conjecture. Although the main theorem of Derksen, van den Essen, and Zhao is stated globally in dimension, its proof has fixed-dimensional content: a noninvertible cubic-homogeneous Keller map in $r$ variables forces the failure of ${\mathrm GMC}(2r)$. Tracking a standard Bass--Connell--Wright reduction of the announced map gives a conservative cubic-homogeneous counterexample in $79$ variables, and hence a route-based failure of ${\mathrm GMC}(158)$. That route is nonconstructive at the final Gaussian step and does not furnish explicit polynomials $P,Q$. The much smaller explicit failures in dimensions $4$ and $3$ below were not derived from the announced Jacobian map.
Let $X_1,\ldots,X_n$ be independent Gaussian tensors in $\mathbb{R}^{d_1}\otimes\cdots\otimes\mathbb{R}^{d_k}$ with a common covariance matrix given by the Kronecker product of $k$ unknown positive-definite factors, and let $D=\prod_{a=1}^k d_a$ and $d_{\max}=\max_a d_a$. Franks et al. (2026) established condition-number-free guarantees for the tensor-normal maximum likelihood estimator under the sample-size condition $nD\gtrsim k^2 d_{\max}^3$ and asked whether the cubic dependence on $d_{\max}$ could be reduced to a quadratic one. We answer this question affirmatively. For $t\geq 1$, if $nD\geq C k^2 d_{\max}^2 t^2$, then with high probability the maximum likelihood estimator exists, is unique, and satisfies $d_{\rm FR}(\widehat\Theta,\Theta)\leq C t \sqrt{k} d_{\max}/\sqrt{n}$ and $d_{\rm FR}(\widehat\Theta_a,\Theta_a)\leq C t\sqrt{k d_a} d_{\max}/\sqrt{nD}$ for every mode $a$. For every mode $a$ with $d_a=d_{\max}$, we further establish the sharp Thompson-metric bound $d_{\rm op}(\widehat\Theta_a,\Theta_a)\leq C t d_{\max}/\sqrt{nD}$. These guarantees are uniform over the unknown covariance factors and require neither condition-number bounds nor sparsity assumptions. Gaussian submodel lower bounds match the full and largest-factor Fisher--Rao rates up to a factor of $\sqrt{k}$ and the largest-factor Thompson rate up to universal constants. Consequently, for fixed $k$, the quadratic dependence of the sample-size threshold on $d_{\max}$ is optimal. GPT-5.6 Sol and Claude Fable 5 were used to assist with proof development, verification, and manuscript preparation.
One implicit DDIM inversion step is the cheapest probe of whether a pretrained diffusion model encodes local manifold geometry. It is the stationarity condition of an explicit potential, $x-G(x)=\nabla\Psi_t(x)$, strongly convex at the Bayes limit with modulus exactly $e^{-h_t}$ for the step's log-SNR gap $h_t$ $-$ for every data law, schedule and point, with no manifold, reach or unimodality hypothesis. Three consequences must be kept apart. (i) The solution is unique at the Bayes limit; a second one requires the trained score to violate the posterior-covariance bound by $1/(1-e^{-h_t})$, a hypothesis-free certificate of model error; the same bound makes contraction a schedule constant, $\rho_g^{\star}=1-e^{-h_t}<0.326$ throughout the standard DDPM schedule. (ii) The solver can still fail: Picard iteration is unit-step gradient descent on $\Psi_t$, unstable wherever $\lambda_{\max}(\nabla^2\Psi_t)>2$, so oscillation certifies nothing; damping below $2/\lambda_{\max}$ cures it. (iii) The geometry lives in the convergence domain: on the scale-free depth $w=r\kappa_{\max}$ the oscillation shell sits at $w=\tfrac12$, schedule-free, and the divergence shell at $w=1/(1+\rho_g^{\star})$, with a measured finite-noise correction in $\|\mathrm{II}\|^2$. Exact scores reproduce both to within $0.54\%$ on three classes; no trained score we probe shows a shell $-$ a derived limitation, not a null result: the Fermi window conflicts with the model's own training support by $3.6$-$5.6\times$, and the trained Hessian-Lipschitz constant is $2$-$12\%$ of the curvature the law reads, $0$ on a ReLU net. Finally the unconditional ceiling $\sigma_t\lambda_{\max}(\mathrm{sym}\,J)\le1$, from $\mathrm{Cov}(x_0\mid x_t)\succeq0$ alone, holds for the exact score to $3\times10^{-7}$ but is violated in all DDPM CIFAR-10/CelebA-HQ-256 settings, by $1.26$-$4.66\times$.
We study the simultaneous approximation of constant-degree polynomials over convex sets. For any family of $m$ degree-$d$ polynomials and any convex set ${H} \subseteq \mathbb{R}_{\ge0}^n$, we construct an $\epsilon$-Cover of the joint value set $\{(f_1(x), \dots, f_m(x)) : x \in {H}\}$ in the $\ell_\infty$-norm. This cover is of size $n^{O(\log(mn)/\epsilon^2)}$, provided the polynomials have constant range over the smallest $\ell_1$-ball inscribing ${H}$. Our approach extends classical net-based sparsifications for linear functions (e.g., Lipton, Markakis, and Mehta [2003]) to arbitrary families of constant-degree polynomials over general convex sets. We use a two-step scheme: first, we construct a quasi-polynomial pre-cover of the family on the smallest $\ell_1$-ball containing ${H}$ by using a concentration argument and leveraging a connection between Bernstein approximation and multinomial distributions; we then compress the pre-cover to ${H}$ by using a recursive degree reduction and feasibility programs anchored at points of the pre-cover. The existence of these covers immediately yields a unified framework for Quasi-Polynomial Time Approximation Schemes (QPTAS) across a wide range of a problems, including fixed-degree polynomial minimization over polyhedral sets, Constraint Satisfaction Problems (CSPs), Free Games, variational inequalities with polynomial operators (which implies guarantees for local Nash equilibria in polynomial games), and additive approximation for normalized densest $k$-subhypergraph on $O(1)$-uniform hypergraphs.
Martino Bernasconi, Matteo Castiglioni, Andrea Celli et al.· 0 citations