There is a number $\psi$ such that for all $\varepsilon>0$ the probability that the value of the output neuron is in $[\psi - \varepsilon, \psi + \varepsilon]$ tends to 1 as $n$ tends to infinity.
Abstract
We consider for an arbitrary fixed $\rho$ and for each positive integer $n$ a multilayer feedforward artificial neural network with $\rho$ layers, $n$ neurons in the first layer (the input layer) and only one neuron, the output neuron, in the last layer. Very roughly formulated, the main result is that if the distribution of weights of connections from a layer to the next are, for all large $n$, approximated well by a fixed continuous (but otherwise arbitrary) curve which does not depend on $n$, and if the values of the $n$ input neurons are independently and identically distributed with a continuous probability density function, then there is a number $\psi$ such that for all $\varepsilon>0$ the probability that the value of the output neuron is in $[\psi - \varepsilon, \psi + \varepsilon]$ tends to 1 as $n$ tends to infinity.
Leveraging the neural architectures which we introduced in arXiv:2109.13512v4, we show a global universal approximation theorem in the topology of $L^p(\mu)$, where $1\le p<\infty$ and $\mu$ is a Radon probability measure on a suitable infinite dimensional topological space $\mathfrak X$. Namely, any function $f:\mathfrak X\to \mathbb R$ in $L^p(\mu)$ can be approximated to any degree of accuracy by suitable infinite dimensional architectures. These architectures can be in turn approximated by almost classical neural networks which are specified by a finite number of parameters only. The vectorial case (where $f=f(x)\in E$ and $E$ is a Banach space) is also considered and analogous results are obtained.
The techniques share most of the high-level ideas presented in [Ruess et al., 2026], but there are also some minor differences which may be of interest for future research on this problem.
We investigate the best $L_2$ approximation of mixed Sobolev spaces by shallow neural networks with $n$ neurons and general activation functions. We first establish an activation-independent Fourier-block principle: if an activation has univariate approximation order $\rho$ in the sense of the Fourier-block property, then the global approximation rate has algebraic order $\min\{\alpha,\rho\}$ for target functions of mixed smoothness $\alpha$, up to explicit logarithmic factors. To verify this property for concrete activations, we introduce a structured univariate approximation condition that implies the Fourier-block property with explicit parameters. For $\mathrm{ReLU}^k$, a matching algebraic lower bound identifies $\min\{\alpha,k+1\}$ as the optimal algebraic approximation exponent in any dimension, up to logarithmic factors in the upper bound. The framework also yields the exponent $\min\{\alpha,k+1\}$ for cardinal B-splines and soft-$\mathrm{ReLU}^k$, and the full mixed-smoothness exponent $\alpha$ for ELU and cosine activations, again up to logarithmic~factors.
Shallow networks with prescribed or randomly sampled hidden parameters are widely used as numerical trial spaces, yet their optimal Sobolev approximation power with standard smooth sigmoidal activations in general dimension remains unresolved. We establish the corresponding optimal rates for a class of smooth sigmoidal activations with Schwartz-class derivative decay, including $\tanh$, the logistic sigmoid, and the error function $erf$. We first construct deterministic direction--offset dictionaries with $M$ features such that every $u\in H^k(\Omega)$ can be approximated with error of order $M^{-(k-m)/d}$ in $H^m(\Omega)$ for all $0\le m\le k$. This rate is optimal in the sense of Kolmogorov widths for Sobolev balls. We further prove that dictionaries obtained by independent parameter sampling from any prescribed density bounded away from zero attain the same approximation exponent with high probability, up to logarithmic oversampling. The analysis develops a sigmoidal ridge representation and combines it with deterministic or probabilistic quadrature in direction--offset space while retaining polynomial control of the output coefficients. Numerical experiments across a broad range of dimensions, target regularities, and Sobolev error norms recover the predicted algebraic rates for both deterministic and random feature dictionaries.
An input may activate few hidden units even when different inputs collectively use an entire network. We study the statistical complexity of this input-dependent sparsity in the one-hidden-layer ReLU model of Awasthi et al. (COLT 2024). For width $s$, at most $k$ active units per input, and effective weight and bias bounds $W,B$, every size-$m$ sample in the class's fixed radius-$R$ input domain satisfies $\mathcal{R}(S)\le CWR\min\{k,\sqrt{sk/m}\log^{3/2}(2m)\}+kB/\sqrt m$. A support-preserving cover and a single normalized chaining argument remove the previous explicit dimension factor, up to logarithms. Lower bounds on appropriate i.i.d. marginals match up to those logarithms, showing how changing active units across inputs retains a width dependence. The input domain matters: zero-bias networks sparse on the entire ball have at most $2k$ nonzero units and complexity $O(kWR/\sqrt m)$, whereas bias bounds comparable to $WR$ restore the worst-case rate on that same domain in only logarithmic dimension. A spherical-cap construction proves the latter claim without assuming sparsity merely on the sampling support. For a specified normalized bounded loss and biases comparable to $WR$, we also obtain agnostic minimax excess-risk bounds of order $\min\{1,\sqrt{s/(km)}\}$ up to logarithms.
Xiaoyu Li, Zhizhou Sha, Jiaojiao Jiang et al.· 0 citations
We establish high-probability bounds for mixed input derivatives of wide random neural networks whose activation derivatives satisfy a factorial growth bound. Our main result specializes these estimates to $\tanh$ networks with Xavier initialization. A direct deterministic analysis based on Euclidean operator norms of the weight matrices yields derivative bounds that generally grow exponentially with the depth. We show that this growth can be substantially improved for sufficiently wide Gaussian networks by isolating the term that is linear in the highest-order derivative and controlling the corresponding tangent directions by measurable finite nets. For scalar-output $\tanh$ networks with Gaussian weights and Xavier initialization, we prove that there exist constants $C,C_0,C_1>0$ such that, whenever the common hidden width satisfies $n \geq C\left(L^3n_0^2(1+\log n_0)+L^2\left(1+\log(L/\eta)\right)\right)$, then, with probability at least $1-\eta$, the estimate $\left|D^u\mathcal{R}_{\Phi^{(L)}}(x)\right| \leq C_0 |u|! (C_1L)^{|u|-1}\prod_{j\in u}\beta_j(\eta,n_0)$ holds simultaneously for every non-empty $u\subseteq[n_0]$ and every $x\in[0,1]^{n_0}$. Thus, the first-order derivative bound is independent of the depth, while a square-free mixed derivative of order $|u|$ grows at most polynomially as $L^{|u|-1}$, apart from the coordinate factors. As consequences, we obtain high-probability bounds for the Euclidean Lipschitz constant and for weighted Sobolev norms of the network realization. The latter connect the derivative estimates to quasi-Monte Carlo integration and indicate how such regularity can enter the analysis of QMC-based training.
Josef Dick, Michael Feischl, Fabian Zehetgruber· 0 citations