TRAP: Understanding and Mitigating Privacy Memorization in Language Models
TRAP, a one-sided penalty on tokenwise TRA that acts only where the target model pulls ahead of its reference, brings memorization near the level of an untrained model at little utility cost, where generic regularizers barely move and differential privacy gives up most of what fine-tuning bought.