Preprint
Aug 2026
Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders
This work applies Top-K Sparse Autoencoders to the intermediate representations of DeepSeek-R1-Distill-Qwen-7B and examines the model's divergent behaviors across math-solving tasks of three distinct difficulty levels, identifying a clear distinction in how the model functions under two reasoning modes.
Bo Cheng, Qiaolin Lu, Yi Chang et al.
· 0 citations