Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed...
Dongwon Jung, H. Ramesh, Yi-Fan Wang et al.· 0 citations
In security-sensitive applications, language-model agents are often required to coordinate without disclosing confidential information. Yet repeated interactions may also let ordinary messages acquire shared private meaning. We study a repeated game with pairs of models in which the sender model observes one of four se...
Jacob Dineen, Si-Lei Ren, Mu-Hao Chen et al.· 0 citations
SafeClawArena is developed, a benchmark of 406 adversarial tasks executed in containerized replicas of real agent platforms with canary-marked credentials and evaluated via automated taint tracking across nine output channels, exposing the inadequacy of current defenses and suggesting directions for future hardening.