Intent-Hiding Jailbreaks: An Information-Theoretic Framework for Compositional Attacks
Recent work has shown that large language models (LLMs) can be vulnerable to jailbreak attacks in which harmful intent is obscured through composition with benign tasks. A harmful request refused in isolation may elicit a different response when embedded within a larger, seemingly benign query. We study these compositi...