Skip to content

Interpretable GOHR Agents via Sparse Autoencoders

Jul 2026 · arXiv.org · Vol abs/2607.25132 · 0 citations · 5 references
Computer Science

TL;DR

This work reports interpretability experiments for a tokenized autoregressive Transformer agent in the Game of Hidden Rules (GOHR), a compact two-rule task in which both hidden rules map object shapes to target buckets, but with different permutations.

Abstract

A central challenge in interpreting learned decision-making systems is to determine whether their internal representations contain concepts that help explain their behavior. We report interpretability experiments for a tokenized autoregressive Transformer agent in the Game of Hidden Rules (GOHR). We focus on a compact two-rule task in which both hidden rules map object shapes to target buckets, but with different permutations. The policy is trained on episodes sampled from these two hidden rules and then evaluated with fixed weights. It is never given a rule label and does not use an explicit rule classifier; any rule information must be inferred implicitly from interaction history. In this setting, the correct rule is not identifiable before the agent tries an informative move and observes accept/reject feedback. Sparse autoencoders (SAEs) trained on the agent's decision-token embeddings recover this structure. When held-out decisions are labeled by simple concepts such as the chosen shape or bucket, SAE dimensions that are highly selective for a concept cover most decisions where that concept is present. Individual SAE dimensions also correspond to interpretable strategies such as probing one rule hypothesis and switching after negative feedback.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Principled Thoughts for Latent Recursive LLM Systems

Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. Theoretical and empirical analyses establ...

Fahd Seddik, F. Fard · 0 citations
#artificial intelligence Preprint Sep 2026

Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It

We study training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and the two engines assign different probabilities to the same tokens. To account for this discrep...

Tian-Run Yu, Kai-Xiang Zhao, Shang-Zhe Li et al. · 0 citations
Preprint Aug 2026

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation

Observation-Calibrated Self-Distillation (OCSD), which contrasts two structurally matched replay views, Full and Observation-Ablated, to derive an observation residual that discounts score changes shared by the replay scaffold, and applies this residual to modulate token-level GRPO updates at high-uncertainty steps, wh...

Yi Yang, Congming Qin, Xiaodan Liu et al. · 4 citations
#natural language process... Preprint Sep 2026

Bongard: Training Machine Intuition

Human intelligence relies heavily on learned intuition: recognising patterns and judging situations without explicitly unfolding every intermediate step. We introduce Bongard, an open-weight System One model that treats machine intuition as an independent capability to design and train. A T5Gemma 2 4B-4B encoder-decode...

Li Ding, Hai-Di Jin, Chen-Hua Ji · 0 citations
#machine learning Preprint Sep 2026

Understanding LLM Parameter Update Sparsity through the Lens of Fisher

Recent studies have observed that parameter changes during language-model post-training can be concentrated in a small subset of coordinates. This phenomenon has been reported in reinforcement learning, on-policy distillation, and supervised fine-tuning on near-policy data. Its recurrence across different post-training...

Yu-Fan Zhang, Sagnik Mukherjee, Hao Peng · 0 citations
#artificial intelligence Preprint Sep 2026

OPSRD: On-Policy Self-Role Distillation

Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transf...

Wei-Jie Ren, Yan-Wen Zhang, Hao Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.