Large language models (LLMs) increasingly power autonomous coding agents such as Codex and Claude Code, yet their training corpora may contain confidential credentials exposed in public repositories or collected from private development artifacts, creating risks of memorization and subsequent leakage. Existing extracti...
Shi-Qian Zhao, Si-Wei Jiang, Xin-Feng Li et al.· 0 citations
Diffusion-based text-to-image (T2I) models are increasingly used for visual content creation, making their generation capability a valuable intellectual property asset. However, this capability is vulnerable to black-box output-based distillation, where an adversary queries the service, collects prompt-image pairs, and...
Zi-Han Wang, Bo-Heng Li, Rui Zhang et al.· 0 citations
Recently, large language model (LLM) agents, such as Codex, Claude Code, and OpenClaw, have become capable of planning and executing long-horizon tasks through repeated tool calls. This capability also creates new opportunities for prompt injection. Existing attacks either place the malicious objective in one explicit...
Shiqian Zhao, Yang-Fan Zhou, Xin-Feng Li et al.· 0 citations
ASCon is proposed, a direction-aware reciprocal \textbf{A}gent--\textbf{S}tep \textbf{Con}textualization model for multiple failure attribution targets that introduces direction-aware graph attention to model execution context, masked step-to-agent attention to construct behavior-aware agent representations, and agent-...
Shuyu Jiang, Yue Ran, Kaiyu Xu et al.· 0 citations
This paper proposes LeakGauge, which probes this response by appending a suffix that gauges leakage behavior and mapping its prefill token probabilities to an attack-risk score, and shows that the risk score is sensitive to an internal leakage-related direction.
Maosen Zhang, Jianshuo Dong, Bo-Han Lu et al.· 1 citation
This work proposes Sampled-BPE, a lightweight token-level auditing pipeline that sample a small subset and train BPE tokenizer to surface polluted tokens, and releases a hierarchical Chinese web token dataset with 660k+ token records, organized as trees to support review and tracing of pollution.
Qingjie Zhang, Ziqi Tang, Jie Zhang et al.· 0 citations
A probe-gated reasoning-based defense is introduced to bridge a knowledge-action gap in agentic LLMs when they are exposed to IPI attacks, and an analysis framework is introduced that identifies natural-language explanations strongly correlated with probe-captured signals.