Skip to content

RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding

Sep 2026 · 0 citations · 11 references
Computer Science

TL;DR

Penn Treebank and a 10k-vocabulary stress test confirm that geodesic decoding transfers across corpora and larger |V|, while hyperbolic curvature helps selectively, and characterize boundary collapse in naive hyperbolic recurrence and show how Mobius stabilization restores trainability.

Abstract

Language models under one million parameters matter for edge deployment, domain adaptation, and reproducible research, yet a two-layer LSTM or Transformer at embedding width d = 128 still spends roughly one third of its capacity on the output matrix W_out in R^(d x |V|). We propose Riemannian Language Models (RiLM), which remove that layer entirely: context unfolds as a trajectory on a Riemannian manifold, and next-token probabilities arise from squared geodesic distance between the current state and vocabulary embeddings. The same embedding map serves input and output -- decoding is geometry. We instantiate the framework on flat R^d (Flat RiLM) and the Poincare ball H^d (HypRiLM) with a shared MLP composition map phi (~290k parameters, d = 128, |V| = 2000). Across five seeds on WikiText-2, HypRiLM reaches 54.2 +/- 0.2 validation perplexity versus 87.6 +/- 0.6 for Flat RiLM; tied and matched LSTM, Transformer, and SSM controls remain at 113-147 PPL on WT-2 -- HypRiLM leads by roughly 2x over the strongest tied recurrent baseline (SSM, 113.0 +/- 3.8). Penn Treebank and a 10k-vocabulary stress test confirm that geodesic decoding transfers across corpora and larger |V|, while hyperbolic curvature helps selectively. We also characterize boundary collapse in naive hyperbolic recurrence and show how Mobius stabilization restores trainability. Claims are scoped to controlled small-model comparisons, not full-vocabulary state of the art.

View source

Similar papers

#machine learning Preprint Sep 2026

LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

Extending LILA to dynamically allocate sparsity budgets via KS-scores yields state-of-the-art generative preservation at moderate compression, while uncovering fundamental single-layer architectural bottlenecks at higher compression regimes.

Sankar Behera, D. Singh, Anshika Agnihotri et al. · 0 citations
#machine learning Preprint Sep 2026

Attention Manifolds: Steering or Blocking Language Models by Editing Learned B-Spline Surfaces

In standard transformer attention, a source token sends the same value vector to every receiver. The query determines \emph{how much} to attend but not \emph{what} to extract. This work introduces \textbf{attention manifolds}: learned 2D B-spline surfaces $S_d(q_d, k_d)$ that modulate each value dimension based on the...

N. Mysore · 0 citations
Preprint Aug 2026

Interpreting Language Model Hidden States at Scale

OmniLens is presented, which applies a single lens family to any model-width activation, whether residual stream, attention, or MLP, and combines two independent scaling techniques, which reproduces key published results at substantially lower cost.

Jordan Pettyjohn, Mansi Sakarvadia, Nathaniel Hudson et al. · 0 citations
Open access Aug 2026

DESS: A Robust Uncertainty Layer for Embedding-Space Models

DESS is introduced, a lightweight uncertainty layer that augments an existing embedding model with a predicted mean vector and an independent per-dimension spread vector that provides a modular, geometry-aware uncertainty layer for embedding-space models, provided its spread is calibrated to local embedding geometry.

Morten Grundetjern, J. Voigt, Per-Arne Andersen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

RGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models

Riemannian Geometry-Sensitive Quantization (RGSQ), which formulates quantization as a reconstruction problem under a unified Fisher-Riemannian metric, enabling standard unimodal PTQ methods to evaluate multimodal quantization error under their original assumptions.

Zhi-Ping Wu, Dong-Dong Ren, Yang Zhou et al. · 0 citations
Preprint Aug 2026

TASSO: TAsk-Specific Subspace Optimization for Continual Learning of Vision-Language Models

TASSO, a new paradigm that efficiently preserves the latent space geometry while ensuring network plasticity, is introduced with two complementary techniques: subspace learning and geometry-aware knowledge distillation.

Chang-Ming Sun, Francesco Barbato, Matteo Caligiuri et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.