Activation steering offers an inference-time defense for vision--language models (VLMs) by modifying intermediate representations without updating backbone parameters. However, protection on benchmark inputs may not persist across alternative expressions of the same harmful request. We investigate this gap using fixed...
Xin-Wei Zhang, Ao-Ting Hu, Hang-Cheng Liu et al.· 0 citations
Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the...
XiaoYu Xu, Zi Liang, Min-Xin Du et al.· 0 citations
Text-to-image diffusion models enable data-efficient"mimicry"attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual. A common countermeasure is to embed imperceptible, low-energy watermarks, yet recent studies show these signatures are br...
Hao-Yang Li, Ruo-Xi Sun, Qing-Qing Ye et al.· 0 citations
Multimodal training relies on image-text corpora collected from external sources, creating opportunities for attackers to poison the data. Stealthy attacks can preserve plausible image-text pairs while concealing the differences used by detectors, so apparently clean data can still redirect the trained model. We theref...
Hao-Yang Li, Ya-Xin Xiao, Lin-Yan Dai et al.· 0 citations
As large language models (LLMs) are increasingly redistributed, adapted, and served behind opaque APIs, model ownership can no longer be established reliably by inspecting model internals or deployment records. This creates a need for behavioral signatures that remain observable through black-box interaction. Yet most...
Zhong-Rui Sun, Jia-Hao Chen, Ou-Bo Ma et al.· 0 citations
SEBA is a sample-efficient framework for black-box adversarial attacks on visual RL agents that significantly reduces cumulative rewards, preserves visual fidelity, and greatly decreases environment interactions compared to prior black-box and white-box methods.
Tai-Ran Huang, Yu-Lin Jin, Jun-Xu Liu et al.· arXiv.org· 0 citations
This work proposes Continuous Agents for Injection Threats via Lifelong Yielding Nexus (CAITLYN), an agent-agnostic defense middleware that matches the detection performance of state-of-the-art defenses at lower token overhead than LLM-as-a-judge baselines.
Zi Liang, XiaoYu Xu, Yanyun Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.