Large language models (LLMs) increasingly power autonomous coding agents such as Codex and Claude Code, yet their training corpora may contain confidential credentials exposed in public repositories or collected from private development artifacts, creating risks of memorization and subsequent leakage. Existing extracti...
Shi-Qian Zhao, Si-Wei Jiang, Xin-Feng Li et al.· 0 citations
It is revealed that contexts with stronger semantic-shift capabilities are more likely to guide models toward recovering harmful meanings and achieving successful jailbreaks, and a black-box context-aware semantic-shift jailbreak framework with Iterative Context Optimization is proposed.
Hujian Zhu, Yi-Hao Huang, Felix Juefei-Xu et al.· 0 citations
Text-to-image (T2I) models can be exploited to produce unsafe images. Existing safety measures, e.g., content moderation or model alignment, can be weakened by adversaries who attempt to restore unsafe generation through model fine-tuning. This paper presents Patronus, a defensive framework that improves T2I models’ re...
Xin-Feng Li, Sheng-Yuan Pang, Jialin Wu et al.· IEEE Transactions on Informa...· 0 citations
AEGIS (Adaptive Ensemble Guard for Injection Shielding) extracts instruction-sensitive projectors to identify malicious instructions and leverages a Unified Multi-Layer Consensus mechanism that aggregates topologically distinct signals across the network depth.
Jia-Hao Chen, Ruiping Yin, Xin-Feng Li et al.· 1 citation· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.