How many dimensions does a language model's computation actually use? The question is ill-posed until one names a functional. Task-weighted charts make it well-posed: low-dimensional coordinate systems fit against a chosen functional of the representation, under the functional's own metric, turning distillation into pl...
Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety. We ask whether sentence-level classification signal lives in the Riemannian geometry of contextual token embeddings, and probe it by extracting per-token pullback metrics from a learned encoder's anal...
S. Konior, Alexandre Quemy, Przemyslaw Klocek et al.· 0 citations
This work asks whether a trained network moves this cloud the way optimal transport would: at the cheapest cost, and along the map that pairs each token with its optimal destination, along the map that pairs each token with its optimal destination.
This work fits this flow's equation of motion as a discrete Langevin model over corpus-mean trajectories of Pythia-160M and Pythia-410M, and scores the predicted steps on held-out tokens.
A token-embedding table holds a hub of short rows near its origin, and it is shown that this cluster biases what nearest-neighbor intrinsic-dimension (ID) estimators report, and it is shown that normalizing the rows instead of removing them gives the same lower reading.
Alexandre Quemy· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.