2 papers indexed here
We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.
Not the right person? Other researchers publish under this name.
When Safety Speaks a Language: A Mechanistic Analysis of Safety-Language Identity Entanglement in LLMs
This work presents a systematic mechanistic analysis of multilingual safety using sparse autoencoder features, sparse interpretable directions in the residual stream associated with harmful and harmless model behavior across three instruction-tuned LLMs, eight languages, and all model layers to qualify the language-uni...