Skip to content
Preprint

Data Attribution of Emergent Misalignment with Persona Features

Aug 2026 · 0 citations · 54 references
Computer Science

Abstract

Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies. We ask where these features come from: which pre-training documents activate them, and whether naturally occurring human-written text suffices to induce EM. Using Sparse Autoencoder (SAE) based model diffing across four open-weight models, we find that features related to jailbreak personas, sarcasm, deception, and manipulation are amplified by misalignment fine-tuning, while safety-relevant and assistant-identity features are suppressed. Steering individual features controls EM in both directions: it induces misalignment rates of up to 62% in aligned models -- exceeding the 35% reached by misalignment fine-tuning itself -- and re-aligns misaligned models to near-baseline misalignment rates. Attributing the causal features to a corpus of one million pre-training web documents retrieves semantically relevant narratives about villainous characters, domination, and harmful agency. However, fine-tuning on these human-written documents does not reliably induce EM, even after reformatting into assistant-style responses, whereas synthetic instruction-response pairs derived from the same content do -- and transfer across model families. Semantic relevance alone is therefore not sufficient: response structure or model-generated phrasing plays an important role in inducing EM.

View source

Similar papers

Preprint Jul 2026

Emergent Misalignment Recruits a Pre-existing Persona Subspace

It is found that narrow fine-tuning recruits a persona structure that is present in the model before the fine-tune exists, and that broad misalignment on questions unrelated to the training data is more broad than mechanical weight superposition and matched diversity jointly account for.

Mohammed Suhail B Nadaf · 0 citations

Reversing Emergent Misalignment Using Simple Self-Distillation

This work induces misalignment by fine-tuning a Qwen2.5-14B-Instruct base model on nar-rowly misaligned data and tests Simple Self-Distillation as a method to recovering alignment in misaligned models.

Adam Banks, J. Huang · 0 citations
Jun 2026

Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment

These results identify optimiser choice as a key factor in EM severity, but show that spectral regularisation can substantially mitigate the effects of EM-prone optimisers.

J. R. Brown, Patrick Leask, Lev McKinney · 0 citations
Preprint Jul 2026

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon?

This work systematically study repeated alignment and misalignment cycles using controlled fine-tuning loops while tracking behavioral performance, and LoRA representations throughout training, finding that both misalignment and realignment are highly sensitive to superficial dataset characteristics.

Abhinav Rao, Liancheng Gong, Bin Hu et al. · 0 citations
Preprint Jul 2026

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift

A systematic evaluation of representative task-adaptation methods shows that task adaptation is not merely a capability-improving step, but an alignment intervention in its own right, motivating multi-dimensional alignment evaluation as a standard component of post-training pipelines.

James Elcock, William F. Shen, Xinchi Qiu et al. · 0 citations
Preprint Jul 2026

Misalignment Has a Personality: A Big Five Account of Emergent Misalignment

Calibrated personality vectors transform an opaque safety phenomenon into a human-legible diagnostic profile by extracting activation directions for character traits from a single binary contrast, which can separate or steer behavior without establishing a calibrated scale.

Hasibur Rahman, Smit Desai · 0 citations