Skip to content

Geometry-Aware Bayesian Parameter-Efficient Fine-Tuning on the Stiefel Manifold via Stein Variational Gradient Descent

Sep 2026 · 0 citations · 32 references
Computer Science

TL;DR

This paper builds on the singular value decomposition factorization of adapters to develop a framework based on Stein variational gradient descent (SVGD), which delivers strong model calibration and attains higher prediction accuracy than SVGD and related uncertainty estimation methods that are formulated in Euclidean space.

Abstract

Several geometry-aware approaches to low-rank adaptation have emerged for parameter-efficient fine-tuning of large pre-trained models. These methods aim to take full advantage of the geometric structure of low-rank manifolds for improving the efficiency in subspace utilization and reducing redundancy by enforcing orthogonality constraints during optimization. The strong empirical results of these techniques have motivated further study into whether predictions from such geometry-based adaptation methods could be overconfident. In this paper, we build on the singular value decomposition factorization of adapters to develop a framework based on Stein variational gradient descent (SVGD). In this formulation, the low-rank matrices are transported along the Stiefel manifold to match the targeted distributions while retaining their crucial geometric structure. Since this geometry-aware SVGD approach provides multiple solutions during inference, it supports uncertainty quantification and produces better-calibrated adapters on the Stiefel manifold. Extensive experiments show that our method delivers strong model calibration and attains higher prediction accuracy than SVGD and related uncertainty estimation methods that are formulated in Euclidean space.

View source

Similar papers

Aug 2026

SPIRA: Sparse Information-Geometric Rank Adaptation for Parameter-Efficient Fine-Tuning of Large Pretrained Models.

Downstream adaptation of large pretrained models (LPMs) via full-parameter fine-tuning is computationally prohibitive. Parameter-efficient fine-tuning (PEFT) methods, such as the widely used Low-Rank Adaptation (LoRA), reduce this cost but still parameterize dense updates over the selected weight matrices. This support...

Zhongyi Wen, Zhikai Zhai, Guo-Min Sun et al. · 0 citations
#machine learning Preprint Sep 2026

Explaining f-Divergence-Based Regularization via Local Curvature and Sharpness-Aware Minimization

Divergence-based regularization and Sharpness-Aware Minimization (SAM) are two prominent approaches for improving generalization in deep learning, both motivated by robustness to perturbations. However, their relationship has remained largely unexplored. Building on classical second-order expansions of $f$-divergences,...

Nour Jamoussi, Marios Kountouris · 0 citations
Preprint Sep 2026

Technical note on: Zero-Training Feature-Space Alignment via Information Geometry

Deep vision models often degrade under distribution shift. Test-time adaptation can improve robustness but typically requires iterative optimization, hyperparameter tuning, and multiple forward-backward passes. We propose Zero-Training Fisher Geometry Alignment (ZFGA), a closed-form method that improves robustness unde...

Behraj Khan, T. Syed, Syed Ahmad Chan Bukhari · 0 citations
#artificial intelligence Preprint Sep 2026

Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods

FITR is proposed, a trust-region-based BO method that replaces lengthscale-based scaling by local pullback-Fisher weights and easily generalizes to non-isotropic surrogates, although the gains are more task-dependent in that setting.

Saksham Kiroriwal, Julius Pfrommer, Jürgen Beyerer · 0 citations
#artificial intelligence Preprint Sep 2026

Riemannian Structure and Optimization for a Class of Low-Parametric Orthogonal Matrices

In this paper, we are concerned with matrices formed by block-diagonal factors interleaved with fixed permutations -- a flexible family of structured matrices. This class has recently drawn interest in deep learning architectures for its balanced expressivity-efficiency trade-off, yet efficient computational strategies...

Aliev S. E. Aliev, Maxim V. Rakhuba · 0 citations
#machine learning Preprint Sep 2026

Nested Inductive Bias Framework for SPD Manifold Learning

This work introduces a Nested Inductive Bias framework that utilizes a two-stage diffeomorphic composition to formally pull back non-Euclidean target geometries onto the SPD manifold, and proposes the Rational Conformal Metric (RCM), designed to establish state-of-the-art geometric robustness against outliers by boundi...

Tushar Das · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.