This paper is the first systematic study of whether prompt-token hidden states in contemporary LLMs exhibit Gromov Hyperbolicity (GH), a distance-based measure of tree-likeness.
Abstract
LLM hidden states are ordinary vectors, but the distances among those vectors may still show hierarchical structure. To our knowledge, this paper is the first systematic study of whether prompt-token hidden states in contemporary LLMs exhibit Gromov Hyperbolicity (GH), a distance-based measure of tree-likeness. Using 818,904 sample-layer measurements from ten open-weight models across MATH500, HumanEval, WinoGrande, and TruthfulQA, we build a GH map over four axes: parameter scale, layer depth, model family, and input domain. The clearest pattern is depth, not scale: middle layers usually form a high-relative-hyperbolicity plateau, while final layers often become substantially more tree-like. Scale effects are weak and non-monotonic, matched 7/8B model families differ strongly, and domains interact with model specialization. These findings make GH useful as a practical diagnostic: it shows where hierarchical distance structure appears, how specialization changes it, and which model-layer-domain comparisons deserve closer analysis.
It is shown that layer-wise generative learning can spontaneously uncover and progressively amplify class-related structure in unlabeled data and improve average clustering can coexist with reduced accessibility for a few difficult class pairs.
Patrick Krauss, Achim Schilling, Andreas K. Maier et al.· 0 citations
Large language models trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood, and representation-level evidence that ICL in open-source LLMs approximates optimal Bayesian prediction over a context-inferred generative mod...
Daniel Balcells, Andrew Lee, Chirag Rastogi et al.· 1 citation· ⚡1
A scale-dependent transition between two ID regimes is found: at low lexical diversity, conditions with fewer unique final words produce higher ID, while at high lexical diversity, this ordering reverses, and conditions with more unique words produce higher ID.
Arwa Osman, Marco Baroni, Iuri Macocco· 0 citations
A principled hierarchical architecture for detecting and classifying critical transitions in agent-based opinion dynamics models is established, establishing a principled hierarchical architecture for detecting and classifying critical transitions in agent-based opinion dynamics models.
This work expresses the known accessibility transition in an equivalent conic form, centered for compact convex targets at the statistical dimension of the polar cone, and introduces Random Mapping Networks (RaMaN), which instantiate the predicted latent dimension using structured Hadamard or seed-regenerated Gaussian...
Andrew Cheng, Ali Eslamian, Jie Cheng et al.· 0 citations
Observing what a large language model computes during autoregressive inference--online and without training probes--remains difficult. We introduce Lingtai, a training-free concept telemetry layer: at each generation step, residual states are projected onto a domain-specific bank of named concept anchors, constructed w...
Jian-Gang Chen· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.