MLLMs are increasingly deployed in user-facing applications, yet they inherit backdoor risks from the pipelines used to construct them: triggers may reside in images, texts, or both. Existing model-level backdoor removal methods, largely designed for conventional classifiers, show limited effectiveness on MLLMs, while...
Jia-Li Wei, Ming Fan, Ming-Kun Zhang et al.· 0 citations
Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an accumulated representa...
Hao-Yu Wang, Wei Zhao, Ye-Di Zhang et al.· 0 citations
An automated transformation pipeline that converts existing safety designs into reusable safety skills and establish a community-driven safety skill library is developed and results demonstrate the potential of stage-specific safety skills as a scalable and composable foundation for building resilient and trustworthy a...
LLM agents are moving from single-prompt use to long task streams in which reusable memory becomes a core capability for terminal, software-engineering, and web tasks. Such memory is useful only when stored experience remains reliable across hundreds of interactions, but two failure modes break that assumption in pract...
Hao-Yu Wang, Guang-Yuan Dong, He Liang et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.