Skip to content

Author

Mohamed Amine Merzouk

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

Efficient Safety Alignment of Language Models via Latent Personality Traits

Latent Personality Alignment (LPA) is introduced, which replaces explicit harm refusal with adversarial training on just 66 harm-agnostic statements drawn from psychometric personality literature, hypothesizing that personality-anchored representations share latent structure with harm avoidance, so adversarially stabilizing them implicitly constrains the subspace exploited by jailbreak attacks.

Mohamed Amine Merzouk, Nolan Smyth, Damiano Fornasiere et al. · 0 citations
Preprint Jul 2026

How Much is Left? LLMs Linearly Encode Their Remaining Output Length

Training minimal-capacity linear probes on frozen hidden states of three open-weight 7-8B models across seven completion-style datasets finds three converging pieces of evidence that LLMs maintain a plan-like internal representation of output length, interpreted as evidence that LLMs maintain a plan-like internal representation of output length.

Mohamed Amine Merzouk, Dmitri Carpov, Mirko Bronzi et al. · 1 citation