Long-horizon agents consume external content, invoke tools, and modify persistent state. Indirect prompt injection can exploit task-specific context, propagate across causally connected stages, and alter a consequential action while the workflow continues; we term this staged prompt injection. We build an automated, fe...
Jing-Kai Liu, Yu-Fei Han, Xiaoting Lyu et al.· 0 citations
Conformal Privacy Auditing is introduced, a distribution-free calibration framework that provides a statistical certificate of re-identification risk for each released document against LLM-empowered adversaries and enables audits of open-source models and proprietary API models in a unified framework.
Shuo Huang, G. Haffari, Xing-Liang Yuan et al.· 0 citations
A reverse-training framework is introduced that weakens the trigger-target association, producing low-ASR backdoor models while preserving clean-input performance and exposing a fundamental attacker-defender asymmetry in existing defense paradigms.
Arham Riaz, Ting Yu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.