Skip to content
Open access

Learning Lévy density via adaptive RKHS regression with bi-level optimization

Dec 2025 · Inverse Problems · Vol 42 · 0 citations · 46 references
Computer Science Mathematics Physics

TL;DR

Numerical experiments demonstrate that the bilevel RKHS method provides a more stable and competitive alternative to classical L-curve and generalized cross-validation strategies and that the adaptive RKHS norm is more accurate and robust than Lρ2- and ℓ2-norms for regularization.

Abstract

We propose a nonparametric method to learn the Lévy density from data consisting of the process’s probability densities. We recast the problem as identifying the kernel of a nonlocal integral operator from discrete or noisy data, which leads to an ill-posed inverse problem. To regularize it, we construct an adaptive reproducing kernel Hilbert space (RKHS) whose kernel is built directly from the data. Under source and spectral decay conditions, we show that the reconstruction error decays with the mesh size at a near-optimal rate. Importantly, we develop a generalized singular value decomposition-based bilevel optimization algorithm to select the regularization parameter, resulting in efficient and robust computation of the regularized estimator. Numerical experiments for several Lévy densities, drift fields and data types (PDE-based densities and sample ensemble-based kernel density estimation reconstructions) demonstrate that our bilevel RKHS method provides a more stable and competitive alternative to classical L-curve and generalized cross-validation strategies and that the adaptive RKHS norm is more accurate and robust than Lρ2- and ℓ2-norms for regularization.

Read PDF

Similar papers

Preprint Jul 2026

Statistical inverse learning and $\ell^1$-regularization

It is proved that membership in the approximation space $k_t$ is equivalent to polynomial decay of the best $n$-term approximation error, which is equivalent to polynomial decay of the best $n$-term approximation error.

Abhishake Rastogi, T. Bubba, T. Helin et al. · 0 citations
Open access Aug 2026

Statistical Learning Theory for Inverse-Probability-Weighted Conditional U-Statistics via Delta Sequences Under Functional Missing-at-Random Models

This paper develops a unified asymptotic theory for inverse-probability-weighted conditional U-statistics of arbitrary fixed order in the presence of missing-at-random responses and infinite-dimensional functional covariates. The target is a conditional higher-order functional generated by a measurable response kernel and evaluated locally on a separable Banach space. Localization is formulated through delta sequences, providing a common framework for kernel, partition, regressogram, orthogonal series, and related smoothing procedures without recourse to finite-dimensional density arguments. For bounded kernels, we establish uniform almost-complete convergence over pseudo-compact functional domains and obtain a sharp decomposition into deterministic localization bias and stochastic fluctuation. The latter is governed by the localized-kernel variance, the envelope of the delta sequence, the metric complexity of the indexing domain, and the small-ball concentration of the functional covariate. Unbounded kernels are treated under explicit weighted moment, truncation, and summability conditions. The feasible theory quantifies the additional perturbation induced by estimating the propensity score and identifies conditions under which this first-stage uncertainty is asymptotically negligible. Pointwise distributional theory is derived through a denominator linearization combined with the Hoeffding decomposition of the centered localized kernel. The Gaussian limit is driven by the first projection, while the higher-order canonical components are shown to be negligible under explicit local-mass, moment, and noncancellation assumptions. This yields oracle-equivalent feasible inference, a consistent first-projection variance estimator, and asymptotically valid studentized confidence intervals. A finite-grid adaptive comparison principle is also developed for data-driven resolution selection. The scope of the theory is illustrated through conditional rank functionals, discrimination with incomplete labels, metric-learning criteria, and functional prediction. Synthetic and semi-synthetic studies based on functional classification, phoneme log-periodograms, and growth trajectories document the finite-sample interaction between covariate-dependent label observation, local information loss, propensity estimation, and inverse-weighting variance.

Salim Bouzebda · 0 citations
Preprint Aug 2026

Adaptive Regularization for Random Features: A Neighboring Early-Stopping Rule with Oracle-Rate Guarantees

This work proposes a neighboring early-stopping rule for adaptive regularization in KRR with random features (KRR-RF), using a grid that is uniform in inverse regularization and compares only adjacent estimators, reducing the number of discrepancy comparisons relative to standard all-pairs Lepskii-type procedures.

Caihong Wang, Zhi-Bo Chen, Yue Wang · 0 citations
Preprint Jul 2026

Residual-Christoffel Sampling for Random Feature Collocation of Linear PDEs

Random feature collocation fixes a randomly generated trial space and determines its coefficients from a linear least-squares system. Stability then depends on whether the sampled residual equations represent the geometry induced by the differential operator. We construct an operator-aware discretization in which the operator-applied features determine both the collocation measure and a coefficient whitening map. The randomized scheme combines a residual-Christoffel density with inverse-density weights, while a deterministic scalar-row alternative maximizes successive regularized log-determinant increments. Conditional on the realized trial space, the sampled whitened interior Gram is a spectral approximation to the reference Gram on the retained residual space, with sample complexity linear in the retained dimension up to a logarithmic factor. For uniformly analytic residual kernels, the associated operator has stretched-exponentially decaying eigenvalues and ridge effective dimension that is polylogarithmic in the inverse ridge scale. Experiments on scalar and vector equations, varied geometries, and one to three spatial dimensions show that residual-space sampling and whitening produce numerically full-rank transformed systems with substantially smaller condition numbers and iteration counts. The deterministic construction attains the lowest errors at the smallest scalar sample sizes. Residual-space geometry therefore yields a principled design for stable strong-form random feature collocation.

Jiale Linghu, Yangshuai Wang · 0 citations
Preprint Aug 2026

Amortized Bandwidth Learning for Kernel Density Estimation under Logarithmic Score

Kernel density estimation converts finite samples into probability densities, but its performance depends critically on bandwidth selection. Classical selectors prescribe the sample-to-bandwidth rule analytically or asymptotically, or solve a new optimization for each sample. An amortized framework is proposed that instead learns this mapping across a distribution of density-estimation tasks by optimizing the logarithmic score. A truncated-and-renormalized bounded-support formulation enables stable learning across heterogeneous tasks, while affine standardization allows a selector trained on a single reference interval to transfer across bounded intervals. Experiments under Gaussian sampling, a multi-family benchmark, and randomized Gaussian-mixture training show that the amortized selector consistently and substantially outperforms Silverman's rule, the Sheather--Jones selector, and least-squares cross-validation, with especially large gains in small and heterogeneous samples. Finite Gaussian mixtures provide a generic training mechanism supported by their $L^1$ approximation property. Selectors trained in this way generalize strongly across different density structures, allowing the same trained selector to be applied directly to finite samples from unknown densities without specifying or fitting a distributional family. This combination of broad applicability and strong empirical performance makes the framework attractive for a wide range of applications in which finite samples or ensembles must be converted into continuous probability densities.

Junying Liang, Hailiang Du · 0 citations
Preprint Aug 2026

Diffusion Quasi-Monte Carlo

We study high-dimensional numerical integration with respect to complex target measures using diffusion-based transport maps and randomized quasi-Monte Carlo (RQMC). Score-based diffusion models induce a deterministic probability flow ODE that transports a simple prior to the target, suggesting a principled way to transform low-discrepancy points on the unit cube into informative samples. We construct a cube-to-target map by composing a Gaussian base transformation (the component-wise inverse Gaussian CDF) with an Euler-discretized probability flow ODE. To retain unbiasedness under transport approximation, we formulate integration as importance sampling (IS) on the cube. Our main result provides verifiable conditions under which the resulting IS integrand satisfies the boundary growth condition, implying an $O(N^{-1+\epsilon})$ RMSE for scrambled nets. We then establish these conditions for diffusion probability-flow transport under mild bounded-derivative assumptions on the learned vector field, explicitly controlling the boundary singularities introduced by the inverse Gaussian CDF. Experiments range from a 2D mixture to 784D images and a 40,960D conditional vorticity-assimilation task; in the latter, blocked scrambled Sobol'sampling reduces the randomization standard deviation of nonlinear accuracy metrics at essentially unchanged online denoising cost. Together, these results give a theoretical and empirical foundation for combining diffusion generative modeling with high-precision RQMC integration.

Jianlong Chen, Yifeng Yu · 0 citations