Efficient Safety Alignment of Language Models via Latent Personality Traits
Latent Personality Alignment (LPA) is introduced, which replaces explicit harm refusal with adversarial training on just 66 harm-agnostic statements drawn from psychometric personality literature, hypothesizing that personality-anchored representations share latent structure with harm avoidance, so adversarially stabilizing them implicitly constrains the subspace exploited by jailbreak attacks.