Penn Treebank and a 10k-vocabulary stress test confirm that geodesic decoding transfers across corpora and larger |V|, while hyperbolic curvature helps selectively, and characterize boundary collapse in naive hyperbolic recurrence and show how Mobius stabilization restores trainability.
Abstract
Language models under one million parameters matter for edge deployment, domain adaptation, and reproducible research, yet a two-layer LSTM or Transformer at embedding width d = 128 still spends roughly one third of its capacity on the output matrix W_out in R^(d x |V|). We propose Riemannian Language Models (RiLM), which remove that layer entirely: context unfolds as a trajectory on a Riemannian manifold, and next-token probabilities arise from squared geodesic distance between the current state and vocabulary embeddings. The same embedding map serves input and output -- decoding is geometry. We instantiate the framework on flat R^d (Flat RiLM) and the Poincare ball H^d (HypRiLM) with a shared MLP composition map phi (~290k parameters, d = 128, |V| = 2000). Across five seeds on WikiText-2, HypRiLM reaches 54.2 +/- 0.2 validation perplexity versus 87.6 +/- 0.6 for Flat RiLM; tied and matched LSTM, Transformer, and SSM controls remain at 113-147 PPL on WT-2 -- HypRiLM leads by roughly 2x over the strongest tied recurrent baseline (SSM, 113.0 +/- 3.8). Penn Treebank and a 10k-vocabulary stress test confirm that geodesic decoding transfers across corpora and larger |V|, while hyperbolic curvature helps selectively. We also characterize boundary collapse in naive hyperbolic recurrence and show how Mobius stabilization restores trainability. Claims are scoped to controlled small-model comparisons, not full-vocabulary state of the art.
Extending LILA to dynamically allocate sparsity budgets via KS-scores yields state-of-the-art generative preservation at moderate compression, while uncovering fundamental single-layer architectural bottlenecks at higher compression regimes.
Sankar Behera, D. Singh, Anshika Agnihotri et al.· 0 citations
In standard transformer attention, a source token sends the same value vector to every receiver. The query determines \emph{how much} to attend but not \emph{what} to extract. This work introduces \textbf{attention manifolds}: learned 2D B-spline surfaces $S_d(q_d, k_d)$ that modulate each value dimension based on the...
OmniLens is presented, which applies a single lens family to any model-width activation, whether residual stream, attention, or MLP, and combines two independent scaling techniques, which reproduces key published results at substantially lower cost.
Jordan Pettyjohn, Mansi Sakarvadia, Nathaniel Hudson et al.· 0 citations
DESS is introduced, a lightweight uncertainty layer that augments an existing embedding model with a predicted mean vector and an independent per-dimension spread vector that provides a modular, geometry-aware uncertainty layer for embedding-space models, provided its spread is calibrated to local embedding geometry.
Morten Grundetjern, J. Voigt, Per-Arne Andersen et al.· KI - Künstliche Intelligenz· 0 citations
Riemannian Geometry-Sensitive Quantization (RGSQ), which formulates quantization as a reconstruction problem under a unified Fisher-Riemannian metric, enabling standard unimodal PTQ methods to evaluate multimodal quantization error under their original assumptions.
Zhi-Ping Wu, Dong-Dong Ren, Yang Zhou et al.· 0 citations
TASSO, a new paradigm that efficiently preserves the latent space geometry while ensuring network plasticity, is introduced with two complementary techniques: subspace learning and geometry-aware knowledge distillation.
Chang-Ming Sun, Francesco Barbato, Matteo Caligiuri et al.· 0 citations