2025
For Better or for Worse, Transformers Seek Patterns for Memorization
This work investigates memorization in transformer-based language models by analyzing their memorization dynamics during training over multiple epochs and finds that memorization is neither a constant accumulation of sequences nor simply dictated by the recency of exposure to these sequences.
Madhur Panwar, Gail Weiss, Navin Goyal et al.
· Neural Information Processin... · 1 citation