Preprint
Aug 2026
Locating and Controlling Implicit Personalization in Large Language Models
It is established that a localized internal activation signal tracks changes in recommendations, and it is shown that removing the internal signal associated with one cue can suppress its influence, often more effectively than asking the model to ignore demographics via prompting.
Yueru Yan, Siqi Wu, Thai Le
· 0 citations