Where Do LLMs Decide to Break the Rules? Mechanistic Localization of Prompt Injection Compliance
When a prompt injection attack succeeds, a Large Language Model (LLM) abandons its assigned system role to comply with an adversarial instruction. While prior work has extensively quantified how often this occurs, we ask a more fundamental question: where inside the network does the model actually decide to break the r...