Preprint
Jul 2026
Beyond the Prompt: Jailbreaking Function-Calling LLMs via Simulated Moderation Traces
It is demonstrated that prompt-level sanitization alone is fundamentally insufficient for defending tool-enabled LLM systems and highlight the urgent need for context-aware validation across schemas, arguments, tool outputs, and accumulated conversation state.
Junlong Liu, Haobo Wang, Weiqi Luo et al.
· 0 citations