Aug 2026· Journal of King Saud University: Computer and Information Sciences· Vol 38· 0 citations· 37 references
TL;DR
This work proposes PRISM-ZO, a projection-robust framework that samples low-dimensional random tangent subspaces and combines symmetric finite differences with median-of-means or Huber aggregation and establishes the unbiasedness of the correctly rescaled projected direction in expectation over the random subspace.
Abstract
Zeroth-order optimization on Riemannian manifolds is relevant when objectives are available only through function evaluations, but conventional tangent estimators can be unstable under non-Gaussian oracle errors. We propose PRISM-ZO, a projection-robust framework that samples low-dimensional random tangent subspaces and combines symmetric finite differences with median-of-means or Huber aggregation. We establish the unbiasedness of the correctly rescaled projected direction in expectation over the random subspace, quantify its conditional projection error, prove an explicit expected bound for convergence to an approximate first-order stationary point under a stated estimator-error model, and give a verifiable condition for detecting projected negative curvature. In sparse PCA experiments with 10 random seeds, PRISM-ZO variants are compared with full-dimensional zeroth-order and first-order power baselines under Gaussian, Student-t and Cauchy perturbations; the results report means, standard deviations and 95% confidence intervals and identify both beneficial and failure regimes. The hard-thresholded sparse update is treated as an empirical proxy rather than a smooth retraction.
We study the computation of static mean-field equilibria on a compact state space by formulating the equilibrium condition as a variational inequality over probability measures. We propose an entropic variant of Korpelevich's extragradient algorithm---the Kullback--Leibler Mirror-Prox method---in which Euclidean projections are replaced by relative-entropy proximal steps. Each half-step is therefore an explicit exponential reweighting of the current measure, implemented on a finite state-space discretization. Under Lasry--Lions monotonicity and continuity assumptions, we prove convergence of mesh-refined ergodic averages and obtain finite-iteration Minty-residual and approximate-equilibrium bounds that jointly quantify iteration and discretization errors. Under strong monotonicity, we derive metric convergence rates for the last, best, and averaged iterates. We also develop a KL-type Tikhonov regularization that selects the equilibrium minimizing relative entropy with respect to a reference measure. The framework applies to potential and nonpotential cost operators and does not require differentiability or convexity of the cost in the individual state.
Erhan Bayraktar, Ibrahim Ekren, L. Vy et al.· 0 citations
Heavy tails weaken high-confidence control for the empirical mean. Geometric median-of-means (MOM) also lacks a threshold that moves toward mean efficiency. We propose \emph{HOMER}, or Huber-of-Means for Efficient and Robust Estimation. HOMER aggregates block means through a radial Huber center. Its canonical and pseudo-Huber forms bound each block score and interpolate between median-like robustness and the empirical mean. We establish a Hilbert-space majority theorem and a MOM-order deviation bound under a finite second moment. Canonical HOMER recovers the sample mean inside its quadratic region. Pseudo-HOMER approaches the sample mean as the threshold grows. It also admits asymptotic linearity and consistent sandwich covariance estimation around the population block-Huber target. Under a finite third moment, fixed finite-dimensional projections support mean inference at the usual parametric rate. This result requires growing block sizes and counts, with block sizes increasing faster. Heavy-tailed simulations show that HOMER remains stable when a minority of block summaries is displaced. On clean Gaussian data, both versions closely approach the empirical mean's efficiency. Finite-block sandwich intervals undercovered, especially for skewed functional data. Further studies show failure when contamination affects most blocks or compromises ordinary within-block means.
We propose a family of random feature maps for scalable kernel machines on low-dimensional subspaces, ie on the Grassmannian manifold. Such representations are useful when data classes or clusters are well described by the span of a few samples. Classical Grassmannian kernels, including the projection and Binet-Cauchy kernels, require full Gram matrices, which leads to prohibitive computational and memory costs for large high-dimensional subspace datasets. We address this limitation using random features based on rank-one projections of subspace projection matrices followed by bounded non-linear transforms, either periodic or binary, to control the resulting distributions. We show that inner products in the random feature space approximate well-defined rotation-invariant Grassmannian kernels that depend only on the principal angles between subspaces. When the number of features is sufficiently large relative to the intrinsic subspace dimension, the approximation holds uniformly over all fixed-dimensional subspaces with high probability. For periodic transforms, the approximated kernel has a closed-form expression with tunable behaviour between inverse Binet-Cauchy and Gaussian-type regimes. Binary transforms yield compact one-bit subspace features, although no closed-form kernel is known. Structured rank-one projections based on randomised fast Fourier transforms further reduce computation without sacrificing practical accuracy. Experiments on synthetic data and ETH-80 classification tasks show that these features accurately preserve Grassmannian geometry while reducing computation, memory, and storage. Rank-one embeddings therefore provide a practical and scalable alternative to classical Grassmannian kernels.
Recovering two-dimensional Ito generators from trajectory data is difficult because drift increments have low signal-to-noise, bivariate weak designs can be ill-conditioned, and unconstrained tensor estimates need not be positive semidefinite. We study WG-SINDy estimator combining covariance-shaped spatial kernels, a ridge-stabilized local-polynomial projection, adaptive-LASSO/STLSQ selection, one in-sample per-component feasible diagonal GLS pass, and a PSD projection--Cholesky read-out with mild isotropic shrinkage. The released estimator uses a data-dependent full-cloud smoother and one in-sample per-component feasible diagonal GLS pass; accordingly, we do not claim exact finite-sample martingale cancellation or a feasible-GLS efficiency theorem for the reported implementation. We evaluate the estimator on 29 synthetic two-dimensional systems: 19 meet their declared per-system recovery contracts, eight are retained as named limits, and two remain scoped reviews. Across the 19 PASS rows, the median central-grid drift metric is 0.204 and the median tensor error is 0.0397. Among the six systems with a finite, non-degenerate off-diagonal target, the median $a_{12}$ cosine is 0.997. Positive-semidefinite validity is imposed by construction. These results are synthetic, in-sample sampled-region diagnostics and do not establish universal or real-data recovery.
We address two open questions in streaming PCA via Oja's algorithm: sharp operator-norm convergence for general rank under sub-Gaussian data, and distributional inference for the resulting subspace estimator. Existing convergence analyses, even in the rank-one case, either assume bounded data or leave non-vanishing remainder terms that prevent adaptation to a polynomially vanishing tail spectrum, while existing distributional results are confined to the rank-one case. Our convergence theory removes these remainder terms and yields a sharp rate. In the dense-tail spiked covariance regime, this rate matches the minimax rate up to logarithmic factors. More generally, we prove a matching lower bound, up to logarithmic factors, across both dense-tail and sparse-tail regimes under a mild nondegeneracy condition. The analysis yields a linearization of Oja's iterates, which in turn enables a high-dimensional Gaussian approximation for the general-rank subspace estimation error with an explicit limiting covariance. We also establish a row-wise Gaussian approximation over convex sets for the aligned difference, recovering prior rank-one results as special cases. For practical inference, we develop an online multiplier bootstrap algorithm and prove its consistency. Beyond streaming PCA, our techniques contribute to Gaussian approximation and bootstrap inference for nonconvex stochastic approximation.
We introduce an intrinsic spectral sparsity model for nonparametric density estimation on compact connected Riemannian manifolds. Instead of penalizing coefficients in an arbitrarily chosen Laplace--Beltrami eigenbasis, we group each complete eigenspace and measure the Hilbert norm of its spectral component. The resulting block-variation space is basis independent and isometry invariant. We establish its structural, atomic, and nonlinear approximation properties and clarify its relation to Sobolev, Besov, and coefficientwise spectral $\ell^1$ classes. We then construct a coordinate-free block-shrinkage estimator and prove a nonasymptotic signal-dependent $L^2$-oracle inequality that adapts to the unknown set of detectable eigenspaces. Under polynomial spectral growth, the risk theory separates the number of spectral blocks from their multiplicities and exhibits two regimes: one driven by a single high-dimensional eigenspace and the other by cumulative spectral complexity. Under matching spectral-growth and nondegeneracy assumptions, corresponding minimax lower bounds show that this multiplicity dependence is intrinsic, with sharp consequences for spheres and the rotation group $SO(3)$. Finally, we develop a positive, normalized, block-penalized exponential spectral sieve for log-densities and derive likelihood oracle inequalities together with expected Kullback--Leibler, Hellinger, and $L^2$ risk bounds. The resulting framework provides a geometry-respecting theory of sparse density estimation that remains invariant under changes of eigenbasis.