LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails
LongGuard is presented, a framework that evaluates, mechanistically analyzes, and mitigates long-context guardrail failure, and proposes two training-free mitigations - Chunked Detection and Attention-Head Sharpening (AHS) - and a deployment protocol that selects configurations by context length and audit side.
Zi-Yang Chen, Xing Wu, Songlin Hu
· 0 citations