Skip to content

Comparing Latent Concept Formation in State Space Models and Transformers via Sparse Autoencoders

Sep 2026 · 0 citations · 14 references
Computer Science

TL;DR

The results suggest that architectural routing mechanisms may have negligible impact on core semantic understanding, with representational divergence confined to extreme structural margins.

Abstract

The quadratic scaling of Transformer self-attention has driven the adoption of sub-quadratic Selective State Space Models (SSMs) like Mamba, which compress past context into a fixed-size recurrent hidden state. This strict informational bottleneck raises a foundational question for mechanistic interpretability: do SSMs and Transformers learn fundamentally distinct latent representations? In this work, we employ Sparse Autoencoders (SAEs) to conduct a large-scale, feature-level correspondence analysis between Mamba-130m and Pythia-70m over a 10-million token corpus. Contrary to hypotheses predicting widespread architectural divergence, we find no evidence of systematic representational divergence between architectures: across the observed Jaccard distribution, 99.98% of Mamba features cluster toward the upper alignment boundary, providing preliminary feature-level support for the Universality Hypothesis. We further identify and qualitatively characterize this microscopic fraction (0.02%) of diverging features, finding patterns consistent with the hypothesis that the recurrent bottleneck selectively limits the parsing of rigid syntax rather than broad semantic ontology. We demonstrate that while Pythia's unconstrained attention permits the monosemantic decomposition of distinct formatting edge-cases, Mamba is forced to compress unrelated syntactical anomalies into polysemantic"junk drawer"neurons to preserve state capacity. Collectively, these results suggest that architectural routing mechanisms may have negligible impact on core semantic understanding, with representational divergence confined to extreme structural margins.

View source

Similar papers

Preprint Aug 2026

Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

It is found that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states, but instead track model-internal similarity structure.

Nikolai Bolik, Lennart Stöpler, Artur Andrzejak · 0 citations
#machine learning Preprint Sep 2026

The Dynamics of Continuous Mixture Collapse in Language Models

This work studies why pretrained language models often fail to preserve mixtures of many components and shows that exact preservation generally requires context-dependent correction, whose required dimensionality can grow with the number of components.

Ali Backour · 0 citations
Open access Sep 2026

Latent Space Representation Learning Based on Variational Autoencoders

The study demonstrates that the VAE possesses significant advantages in state compression and uncertainty modeling, but suffers from issues such as blurry image generation and oversimplified posterior distribution assumptions.

Xiao-Zhou Gao · 0 citations
#machine learning Preprint Sep 2026

A Comparative Analysis of Attention versus State-Space Models for In-Context Learning

Transformers and state-space models (SSMs) are two prominent sequential learning architectures, yet their comparison remains largely empirical and existing theoretical analyses are typically task-specific or architecturally restricted. In this paper, we develop belief geometry, a unified analytical framework for compar...

Enes Arda, Semih Cayci, Atilla Eryilmaz · 0 citations
#artificial intelligence Preprint Sep 2026

Does Joint-Embedding Predictive Architecture Pretraining Help Time Series Forecasting?

Joint-embedding predictive architectures (JEPA) have emerged as a promising self-supervised pretraining paradigm for time series, learning representations by predicting target embeddings in latent space rather than reconstructing raw signals. Yet evidence on their benefits remains mixed, and most studies test only a si...

Yu-Tong Feng, Bo-Wen Liao, See-Kiong Ng et al. · 0 citations
#machine learning Preprint Oct 2026

Learning Rate Transfer for Hybrid Transformer-SSM Architectures

We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Model (SSM) blocks, a class adopted by several recent production language models. In particular, we focus on the gap between the theoretical scaling rules derived for SSMs under zero-order-hold (ZOH) discretization at inf...

Jimin Seo, Gyubok Lee, Yeonsik Jo et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.