FakeMark is presented, a gradient-guided false-claim attack for image classifiers that uses a white-box surrogate but never queries or accesses the victim model during attack construction, to motivate provenance-aware, multi-factor ownership protocols.
Yu-Tong Wu, Wen-Yue Li, He-Wang Nie et al.· Cybersecurity· 0 citations
Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignment. Because deployed LLMs are commonly exposed only through query APIs, ownership verification must often rely on black-box text responses. This setting is difficult: generations are...
Yu-Tong Wu, Xiao-Fan Bai, Shixin Li et al.· 1 citation
This paper introduces the concept of instruction-dense visual jailbreaks, in which image-generation models produce detailed, readable, and actionable harmful instructions within images, and proposes TYPO, a black-box framework that exploits this safety gap by automatically generating adversarial TYPOgraphy prompts.
Meng Xie, Li Zeng, Hang Zhang et al.· arXiv.org· 0 citations
CloakDiff is proposed, the first framework for reversible, high fidelity privacy protection against text-based query attacks in VLMs and EDM Heuristic Sampling, a principled diffusion schedule for adversarial guidance.
Qinghua Lu, Ziqi Zhou, Yufei Song et al.· 0 citations
The proposed LLM agent early-stopping cascade outperforms the best single-gate baseline in every model-environment pair, saving 1.5-8.8 times more compute at a 90% recall target.
Kai Ruan, Zihe Huang, Ziqi Zhou et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.