Skip to content

Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers

Sep 2026 · 0 citations
Mathematics Computer Science

TL;DR

These results identify conditions under which the resulting predictor exploits local geometry and attains the aggregate minimax rate, and derive an in-context generalization bound for near empirical risk minimizers over this class.

Abstract

Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existing nonparametric ICL theory has largely focused on Euclidean domains or single-manifold models. To address this gap, we study the prediction problem under unknown local geometry, modeled by sample size-dependent mixtures of manifolds with heterogeneous dimensions, smoothness, and sampling masses. Under local separation and small-perturbation conditions, we establish a minimax lower bound capturing the aggregate difficulty of the components and construct an oracle tangent local-polynomial estimator with a matching upper bound. This estimator is connected to a structure-informed, two-stage softmax transformer with a geometric preconditioner and chartwise reduced local-polynomial solvers. The transformer achieves negligible approximation error relative to the minimax rate with logarithmic depth and polynomial size. Finally, we derive an in-context generalization bound for near empirical risk minimizers over this class. Together, these results identify conditions under which the resulting predictor exploits local geometry and attains the aggregate minimax rate.

View source

Similar papers

#machine learning Preprint Aug 2026

Sample-Weighted End-to-End Trace-Norm Geometry for Multitask Learning

Multitask models combine a shared representation with task-specific outputs, but generalization bounds often control the two components separately. Such products can discard relative orientation and cancellation and can change under equivalent transformations of intermediate coordinates even when the represented predic...

Mahdi Mohammadigohari · 0 citations
#machine learning Preprint Sep 2026

LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

Extending LILA to dynamically allocate sparsity budgets via KS-scores yields state-of-the-art generative preservation at moderate compression, while uncovering fundamental single-layer architectural bottlenecks at higher compression regimes.

Sankar Behera, D. Singh, Anshika Agnihotri et al. · 0 citations
Preprint Aug 2026

Finite-Sample Metric Non-Collapse for Geometrically Supervised Latent World Models in Control

A finite-sample learning-to-control theory for geometrically supervised latent models of nonlinear deterministic systems is established and an encoder-only local--global metric hinge is introduced that enforces directional resolution and separated-state discrimination.

A. Bensoussan, M. Phung, Minh-Binh Tran · 0 citations
#machine learning Preprint Sep 2026

Geometry-Aware Bayesian Parameter-Efficient Fine-Tuning on the Stiefel Manifold via Stein Variational Gradient Descent

This paper builds on the singular value decomposition factorization of adapters to develop a framework based on Stein variational gradient descent (SVGD), which delivers strong model calibration and attains higher prediction accuracy than SVGD and related uncertainty estimation methods that are formulated in Euclidean...

Quang-Duy Tran, Trung Le, Bao Duong et al. · 0 citations
Preprint Aug 2026

Exact Algebraic Computation of Learning Coefficients for Two-Dimensional Singular Models

This work presents the first deterministic algorithm that computes local RLCTs exactly for any two-dimensional model whose Kullback-Leibler distance is contact equivalent to a polynomial, derive a bound on its complexity, and demonstrate its effectiveness for a broad class of models, with applications including polynom...

Grégoire Sergeant-Perthuis, E. Tsigaridas, Jules Tsukahara Cqsb et al. · 0 citations
#machine learning Preprint Sep 2026

Grokking through the Lens of Minimum-Norm Interpolation

Grokking shows that fitting the training data and learning the underlying signal can occur at very different stages. However, existing theories offer limited quantitative insight into how this delayed generalization depends on inductive bias and signal structure. Our work addresses the gap by developing a statistical t...

Gil Kur, Ileana Rugina, C. Dominé et al. · 0 citations

Related blog posts

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

GPT-Lab Aug 28, 2026

We built an AI factory for HVAC control

What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.