This work constructs finite-dimensional de Rham subcomplexes generated by fixed-neuron shallow ReLU shallow ReLU neural networks and proves exactness in arbitrary dimension and provides a geometric sufficient condition for the required linear independence.
Abstract
We construct finite-dimensional de Rham subcomplexes generated by fixed-neuron shallow ReLU$^k$ neural networks, a class of spaces known to provide optimal approximation rates. For neurons of the form $s_i(x)=\omega_i\cdot x+b_i$, we introduce spaces of neural differential forms: differential $p$-forms whose coefficients are the ReLU$^k$ ridge functions $\sigma_{k-p}(s_i)$. These spaces are compatible with the exterior derivative because differentiating a ReLU power lowers its order by one, and for each fixed neuron, differentiation amounts to exterior multiplication by the fixed one-form $ d s_i$. Under a linear independence assumption on the lowest-order family $\{\sigma_{k-d}(s_i)\}_{i=1}^n$, the global complex decomposes into independent neuron-wise Koszul complexes. We prove exactness in arbitrary dimension and provide a geometric sufficient condition for the required linear independence. Numerical experiments based on the resulting complex provide evidence of stable discretizations and of convergence rates consistent with the underlying approximation theory, and exhibit no spurious modes in eigenvalue problems considered.
A spherical-cap construction proves the latter claim without assuming sparsity merely on the sampling support without assuming sparsity merely on the sampling support, and obtains agnostic minimax excess-risk bounds of order up to logarithms.
Xiao-Yu Li, Zhizhou Sha, Jiao-Jiao Jiang et al.· 0 citations
It is found numerically that allowing $\log \Psi$ to be a nonlinear function of the raw multi-valued spin variable is a natural categorical generalization: it preserves the labelling freedom of the local basis and reproduces the one-hot model with strictly fewer parameters, often with improved trainability.
Leveraging the neural architectures which we introduced in arXiv:2109.13512v4, we show a global universal approximation theorem in the topology of $L^p(\mu)$, where $1\le p<\infty$ and $\mu$ is a Radon probability measure on a suitable infinite dimensional topological space $\mathfrak X$. Namely, any function $f:\mathf...
We investigate the best $L_2$ approximation of mixed Sobolev spaces by shallow neural networks with $n$ neurons and general activation functions. We first establish an activation-independent Fourier-block principle: if an activation has univariate approximation order $\rho$ in the sense of the Fourier-block property, t...
We present the first method for native, continuous gradient descent for machine learning models with $p$-adic parameters. Existing native optimizers are discrete, mostly combinatorial searches, as the $p$-adic numbers $\mathbb{Q}_p$ are totally disconnected, with standard losses that are flat away from their minima. To...
Julian Salazar, D. Kanevsky, Matt Harvey et al.· 0 citations
Consider the compressed shift $T=P_{K_\theta}M_z|_{K_\theta}$ on $K_\theta=H^2\ominus\theta H^2$ associated with the atomic singular inner function $\theta(z)=\exp(-a(\zeta+z)/(\zeta-z))$, $a>0$ and $|\zeta|=1$. Let $k_0=P_{K_\theta}1$ and define real powers $T^s$ using a holomorphic logarithm near the singleton spectr...
I. Krishtal, Javad Mashreghi· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.