Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning
Self-patching identifies activation locations where relocating representations substantially improves failed generalization cases and design a simple heuristic strategy which recovers 58--75\% of the oracle headroom in generalization failure.