Skip to content
Preprint

SIR: Self-improving Red-teaming for Compute Use Agents

Aug 2026 · 0 citations · 39 references
Computer Science

Abstract

Computer-use agents (CUAs) are agents powered by vision-language models (VLMs) that perceive a screen and operate an operating system through mouse, keyboard, and terminal interactions to automate everyday digital tasks. Their exposure to untrusted content creates a risk of indirect prompt injection (IPI), where an adversary embeds instructions in content the agent reads to redirect it toward actions that violate the user's intent. Evaluations based on fixed, hand-written injections may underestimate the risk posed by adaptive adversaries. We present SIR, a black-box self-improving IPI framework that (i) composes task-specific injections from a small library of reusable red-teaming principles stated in plain language and (ii) uses an iterative feedback loop to analyze unsuccessful attack trajectories and distill new, named principles that are retained in a shared library and reused across tasks. We target operating-system-level compromise and evaluate outcomes through deterministic checks on filesystem, service, and permission state rather than an LLM judge. An attack counts as successful only when both the adversarial objective and the benign user task are completed in the same execution. We evaluate three frontier CUAs, allowing SIR up to 10 attack attempts per case. It achieves joint attack success rates of 24% on Claude Opus 4.8 and 28% on Gemini 3.5 Flash, compared with 4% and 0%, respectively, for the benchmark's fixed, hand-written injection. The red-teaming principles discovered against one victim also improve attacks against other victims, including a different model family, without additional feedback.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.