Preprint
Jul 2026
Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning
Self-patching identifies activation locations where relocating representations substantially improves failed generalization cases and design a simple heuristic strategy which recovers 58--75\% of the oracle headroom in generalization failure.
Lu Dai, Ziyang Rao, Yili Wang et al.
· 1 citation