Learning Interpretable Code Explanations of LLM Behavior
This work proposes using reinforcement learning to synthesize human-readable Python programs that replicate an LLM’s input–output behavior, providing behavioral rather than internally faithful explanations.
J. Tey, Nick Jiang
· 0 citations