Skip to content
Preprint

Landscape analysis for shallow neural networks: Complete classification of critical points for cubic activation and affine target functions

Jul 2026 · 0 citations · 51 references
Mathematics

TL;DR

This paper shows that the infimum of the loss is always zero and achievable with at least $d$ active and visible hidden neurons -- that is, hidden neurons with non-zero inner and outer weights -- with pairwise distinct pivots, and provides for arbitrary activation degree $d$ a sharp existence/non-existence criterion for global minimizers with necessary structural conditions.

Abstract

In this paper, we study the optimization landscape induced by the true loss for shallow polynomial neural networks (PNNs) with $\mathfrak{h} \in \mathbb{N}$ neurons on the hidden layer, one-dimensional input and output layers, and a monomial activation of degree $d \in \mathbb{N}$, trained against a non-constant affine linear target function. Our first main result provides for arbitrary activation degree $d$ a sharp existence/non-existence criterion for \emph{global minimizers} with necessary structural conditions. We show that the infimum of the loss is always zero and achievable with at least $d$ active and visible hidden neurons -- that is, hidden neurons with non-zero inner and outer weights -- with pairwise distinct pivots. In contrast, if $\mathfrak{h}<d$, then the infimum cannot be attained and any minimizing sequence of parameters necessarily diverges to infinity. In the second main result, we provide a complete classification of all critical points of the loss function for the cubic activation. We show that the loss landscape admits no \emph{local maximizers}, critical points cannot have exactly two distinct pivots, global minimizers require at least three distinct pivots, critical points with no active hidden neurons correspond to \emph{saddle points} only, and consequently, \emph{non-global local minimizers} and non-trivial saddle points arise only in networks where all pivots coincide. Moreover, non-global local minimizers require all hidden neurons to be active and visible with exactly one hidden neuron having a slope sign matching that of the target function. Our second main result also guarantees that each hidden neuron of a critical point that is not a global minimizer has either input-dependent or zero contribution, but has no nonzero input-independent contribution, to its corresponding realization function.

View source

Similar papers

Preprint Aug 2026

Convex Networks Remain Hard to Certify: Dimension-Accuracy Barriers for Lipschitz Constants

The lifted-selector reduction has an inverse-polynomial radial gap, proved through a quantitative theorem for rational cyclic zonogons, and the results apply to Euclidean zonotope radius and positive-semidefinite binary quadratic maximization parameterized by rank.

Pahan Dewasurendra, Subhashini Jayawardhana · 2 citations
Preprint Aug 2026

Representing MAX functions using two-hidden-layer ReLU networks

The techniques share most of the high-level ideas presented in [Ruess et al., 2026], but there are also some minor differences which may be of interest for future research on this problem.

Zhi-Mao Wang, A. Basu · 0 citations
Jul 2026

Shallower ReLU Network Representations via Exact Linear Algebra

It is proved that $\max_n(x)$ is exactly representable with two hidden layers for every $n\leq 12$, and these results improve upon [Bakaev, Brunck, Hertrich, Stade, Yehudayoff, STOC'26], who proved analogous logarithmic bounds with base three.

Kilian Ruess, G. Averkov, Florestan Brunck et al. · 2 citations · ⚡1
#machine learning Preprint Sep 2026

Nearly Tight Rademacher Bounds for Sparsely Activated Neural Networks

An input may activate few hidden units even when different inputs collectively use an entire network. We study the statistical complexity of this input-dependent sparsity in the one-hidden-layer ReLU model of Awasthi et al. (COLT 2024). For width $s$, at most $k$ active units per input, and effective weight and bias bounds $W,B$, every size-$m$ sample in the class's fixed radius-$R$ input domain satisfies $\mathcal{R}(S)\le CWR\min\{k,\sqrt{sk/m}\log^{3/2}(2m)\}+kB/\sqrt m$. A support-preserving cover and a single normalized chaining argument remove the previous explicit dimension factor, up to logarithms. Lower bounds on appropriate i.i.d. marginals match up to those logarithms, showing how changing active units across inputs retains a width dependence. The input domain matters: zero-bias networks sparse on the entire ball have at most $2k$ nonzero units and complexity $O(kWR/\sqrt m)$, whereas bias bounds comparable to $WR$ restore the worst-case rate on that same domain in only logarithmic dimension. A spherical-cap construction proves the latter claim without assuming sparsity merely on the sampling support. For a specified normalized bounded loss and biases comparable to $WR$, we also obtain agnostic minimax excess-risk bounds of order $\min\{1,\sqrt{s/(km)}\}$ up to logarithms.

Xiaoyu Li, Zhizhou Sha, Jiaojiao Jiang et al. · 0 citations
Preprint Aug 2026

Sharp Sobolev Approximation on General Domains by Linearized Shallow Networks with Analytic Activations

We study Sobolev approximation on bounded domains by linearized shallow neural networks whose inner parameters are prescribed independently of the target function. Our main step is a one-dimensional construction for analytic activations. We prove that quasi-Chebyshev parameter sets with univariate resolution $m$ generate fixed feature spaces attaining the sharp $H^r$-to-$H^s$ approximation order $m^{-(r-s)}$ for a class of analytic activations satisfying a quantitative non-cancellation condition on their Taylor coefficients. Combining this result with the ridge-function lifting theorem in [SIAM J. Math. Anal. 30 (1998), pp. 155-189] and its extension to arbitrary quasi-uniform direction sets established in this work, we construct tensor-product-type parameter sets that attain the sharp rate $$\|f-f_n\|_{L^2(\Omega)}\lesssim n^{-\frac rd}\|f\|_{H^r(\Omega)},\quad f\in H^r(\Omega)$$ for all $r>0$. In contrast to the finite-difference construction in [Neural Comput. 8 (1996), pp. 164-177], whose explicit admissibility condition may require an extremely small parameter scale, the proposed parameter sets remain distributed over fixed intervals and are therefore more amenable to practical computation.

Jia Li, Tong Mao, Jinchao Xu · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.