Jul 2026
Interpretable GOHR Agents via Sparse Autoencoders
This work reports interpretability experiments for a tokenized autoregressive Transformer agent in the Game of Hidden Rules (GOHR), a compact two-rule task in which both hidden rules map object shapes to target buckets, but with different permutations.
Shiwei Tan, Yu-Song Zhao, Weiyi Qin et al.
· arXiv.org · 0 citations