Preprint
Jul 2026
Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs
It is shown that finetuning on narrow, factually-defensible, moderation-passing data can cause broad ideological shifts across unrelated domains, while preserving general capabilities, and proposes a methodology to measure two properties: breadth, how far the shift reaches across topics absent from training, and amplification, how much finetuning intensifies the shift relative to few-shot prompting.
Robert Graham, Edward Stevinson, Yariv Barsheshat
· 0 citations