Preprint
Jul 2026
Efficient Safety Alignment of Language Models via Latent Personality Traits
Latent Personality Alignment (LPA) is introduced, which replaces explicit harm refusal with adversarial training on just 66 harm-agnostic statements drawn from psychometric personality literature, hypothesizing that personality-anchored representations share latent structure with harm avoidance, so adversarially stabilizing them implicitly constrains the subspace exploited by jailbreak attacks.
Mohamed Amine Merzouk, Nolan Smyth, Damiano Fornasiere et al.
· 0 citations