Semantic Isomorphism Attacks and Defense Evaluation for Jailbreaking Large Language Models
The safety alignment of large language models (LLMs) faces persistent challenges from jailbreak attacks. While existing methods mostly leverage prompt engineering or adversarial optimization, we identify and formalize an underexplored semantic isomorphism vulnerability where harmful and safe scenarios share highly cons...