When Can We Trust the Sparse Lens? A Certification Framework for SAE Faithfulness
A post-hoc certification framework forparse autoencoders that provides a practical way to determine whether an SAE representation preserves enough of a frozen LM's predictive behavior to support certification of the original model.