Skip to content
Preprint

Operator-Theoretic Generalization Bounds for Multitask Deep Learning

Aug 2026 · 0 citations · 27 references
Computer Science

TL;DR

Operator-theoretic generalization bounds for deep multi-output function classes are developed by representing network layers as Koopman composition operators on vector-valued reproducing kernel Hilbert spaces and derive Rademacher complexity bounds for invertible and width-expanding injective architectures.

Abstract

We develop operator-theoretic generalization bounds for deep multi-output function classes by representing network layers as Koopman composition operators on vector-valued reproducing kernel Hilbert spaces. In vector-valued Sobolev RKHSs, we derive Rademacher complexity bounds for invertible and width-expanding injective architectures. The estimates separate the output-coupling contribution, represented by the trace of the task matrix, from the layerwise operator norms, Sobolev symbol ratios, determinant factors, and restriction constants generated by the linear maps. We then analyze a distinct one-dimensional Brownian/Cameron--Martin regime. Using the exact anchored derivative-norm characterization of the vector-valued Brownian RKHS, we obtain layerwise bounds for domain-preserving scalar linear maps and anchored diffeomorphic activations; the corresponding factors scale as $|W_l|^{1/2}$ and $\|\sigma_l'\|_\infty^{1/2}$, respectively, and do not involve Sobolev smoothness exponents. Because the Sobolev and Brownian results concern different hypothesis spaces, neither is asserted to dominate the other uniformly. We additionally formulate shared operator learning across tasks, prove a finite-rank representer theorem, derive the exact finite-dimensional problem for squared loss, and establish a target-transfer bound when the learned operator is obtained independently of the target sample. Synthetic and MNIST studies examine stabilized Sobolev-inspired and Brownian-inspired complexity proxies; these empirical proxies are not evaluations of the proved bounds for rank-deficient architectures.

View source

Similar papers

Preprint Aug 2026

Sharp Sobolev Approximation on General Domains by Linearized Shallow Networks with Analytic Activations

We study Sobolev approximation on bounded domains by linearized shallow neural networks whose inner parameters are prescribed independently of the target function. Our main step is a one-dimensional construction for analytic activations. We prove that quasi-Chebyshev parameter sets with univariate resolution $m$ generate fixed feature spaces attaining the sharp $H^r$-to-$H^s$ approximation order $m^{-(r-s)}$ for a class of analytic activations satisfying a quantitative non-cancellation condition on their Taylor coefficients. Combining this result with the ridge-function lifting theorem in [SIAM J. Math. Anal. 30 (1998), pp. 155-189] and its extension to arbitrary quasi-uniform direction sets established in this work, we construct tensor-product-type parameter sets that attain the sharp rate $$\|f-f_n\|_{L^2(\Omega)}\lesssim n^{-\frac rd}\|f\|_{H^r(\Omega)},\quad f\in H^r(\Omega)$$ for all $r>0$. In contrast to the finite-difference construction in [Neural Comput. 8 (1996), pp. 164-177], whose explicit admissibility condition may require an extremely small parameter scale, the proposed parameter sets remain distributed over fixed intervals and are therefore more amenable to practical computation.

Jia Li, Tong Mao, Jinchao Xu · 0 citations
Preprint Jul 2026

Neural and Spectral Operator Surrogates on Gaussian Spaces

We prove expression rate bounds of finite-parametric, spectral and neural surrogates for holomorphic maps between separable Hilbert spaces. The surrogates have an encoder-approximator-decoder architecture, with Karhunen-Lo\'{e}ve encoders and frame decoders. We prove expression rate bounds for two classes of finite-parametric surrogates: i) spectral surrogates obtained by N-term truncations of Wiener polynomial chaos expansions and ii) neural surrogates obtained by approximation of parametric maps with deep feedforward neural networks, ReLU and RePU activation functions and uniformly bounded weights. We work under an algebraic decay assumption on the eigenvalues of the covariance of the Gaussian measure on the input space. We obtain convergence rates for mean-square errors, and additionally in first-order Gaussian Sobolev spaces, to account for errors in the approximation of gradients.

C. Marcati, Mario Mari'c, Christoph Schwab et al. · 0 citations
Preprint Aug 2026

Geometry-Constrained Kolmogorov-Arnold Networks: Learning Edge Geometry via Banach Duality

Kolmogorov-Arnold Networks (KANs) replace fixed activations in deep architectures with learnable univariate edge functions, making the choice of edge parametrisation central. Existing variants rely on fixed bases such as splines, polynomials, or Fourier features, which impose a function-space geometry before data are observed. We introduce geometry-constrained KANs, a family of edge activations derived from Banach duality maps in which the geometry itself is learned through a scalar exponent $p>1$ per edge. This exponent controls the qualitative response: sub-Euclidean values produce sharp, threshold-like behaviour reminiscent of the $\ell_1$ (LASSO) geometry, $p = 2$ recovers the linear regime, and larger values produce flatter responses near the origin. Across 50 symbolic-regression targets ($40$ from the AI Feynman benchmark plus $10$ synthetic stress tests), geometry-constrained KANs match or beat every fixed-basis baseline on median NRMSE (Banach-KAN $0.030$, tying Chebyshev and improving on splines); on average rank Banach-KAN is best on the $18$-equation core ($2.00$) and statistically tied with the strongest spline on the full benchmark ($2.32$ vs. $2.34$). The clearest gains appear under measurement noise: as $\sigma$ grows from $0$ to $1$, $\ell^p$-KAN degrades only $3.7\times$ -- below even a cross-validated spline ($\approx 11\times$) -- while an unregularised spline degrades $21.6\times$; Banach-KAN degrades $8.8\times$, comparable to a tuned spline but far more stable than the unregularised one. Banach-KAN also takes the most per-equation wins in the small-sample regime, with fixed-basis models catching up only as the training set grows. Learned exponents provide an interpretable, relative signal: at a fixed initialisation they reveal a consistent, target-dependent geometric ordering across equation families and input dimensions.

K. S. Sesh Kumar · 0 citations
Open access Aug 2026

Synthetic Data-Guided Symmetric Neural Network Approximation in Banach Spaces

This paper develops a Banach space-valued approximation framework based on symmetrized neural network (SNN) operators generated by a deformation-dependent sigmoidal activation function. Symmetry is introduced directly at the activation level through a reciprocal-deformation mechanism, yielding a positive, even, normalized, and localized density kernel satisfying the partition of unity. The resulting construction provides normalized compact-interval and whole-line quasi-interpolation operators for Banach space-valued functions. Quantitative pointwise and uniform convergence estimates are established through the first modulus of continuity and are extended to higher-order and Caputo–Bochner fractional approximation. Numerical diagnostics support the theoretical kernel properties, and fractional approximation experiments compare the SNN and classical NN operators under common computational conditions. A controlled blind-prediction experiment on a synthetic monthly temperature-like series uses a strict fit–validation–test protocol and a parameter-matched operator comparison, with seasonal ARIMA and MLP models as external baselines. Across five independent realizations, the SNN attains the best mean predictive performance, with R2=0.9500, NMAE =0.0452, and NRMSE =0.0570. A vector-valued experiment in Y=R2 further illustrates the non-scalar applicability of the Banach space framework. In addition, the normalized SNN kernel weights provide an intrinsic node-level interpretation mechanism without requiring an external post hoc explainability method.

G. Anastassiou, Seda Karateke, M. Zontul · 0 citations
Preprint Aug 2026

Kernel Methods for Learning Operators with Multiple Inputs and Outputs

This work introduces a general kernel-based encoder-decoder framework for operator learning that separates observation, representation, learning, and reconstruction, and develops this framework for multi-input, multi-output operator learning, where operators map between products of potentially distinct function spaces.

Adrien Weihs, Chunyang Liao, Jingmin Sun et al. · 0 citations
Preprint Aug 2026

Density Estimation on Compact Manifolds under Intrinsic Spectral Block Variation

We introduce an intrinsic spectral sparsity model for nonparametric density estimation on compact connected Riemannian manifolds. Instead of penalizing coefficients in an arbitrarily chosen Laplace--Beltrami eigenbasis, we group each complete eigenspace and measure the Hilbert norm of its spectral component. The resulting block-variation space is basis independent and isometry invariant. We establish its structural, atomic, and nonlinear approximation properties and clarify its relation to Sobolev, Besov, and coefficientwise spectral $\ell^1$ classes. We then construct a coordinate-free block-shrinkage estimator and prove a nonasymptotic signal-dependent $L^2$-oracle inequality that adapts to the unknown set of detectable eigenspaces. Under polynomial spectral growth, the risk theory separates the number of spectral blocks from their multiplicities and exhibits two regimes: one driven by a single high-dimensional eigenspace and the other by cumulative spectral complexity. Under matching spectral-growth and nondegeneracy assumptions, corresponding minimax lower bounds show that this multiplicity dependence is intrinsic, with sharp consequences for spheres and the rotation group $SO(3)$. Finally, we develop a positive, normalized, block-penalized exponential spectral sieve for log-densities and derive likelihood oracle inequalities together with expected Kullback--Leibler, Hellinger, and $L^2$ risk bounds. The resulting framework provides a geometry-respecting theory of sparse density estimation that remains invariant under changes of eigenbasis.

Olga Klopp, Fedor Noskov · 0 citations