Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences the benchmark's measur...
Neeraj Karamchandani, Piyush Nagasubramaniam, Xin-Hong Xie et al.· 0 citations
This work presents a systematic study of honeypot-aware budget allocation for LLM attack agents and shows that with the proposed detector-guided policy, LLM agent attackers can effectively allocate budget to compromise hosts in a host pool, highlighting the importance of dynamically allocating budget in a controlled mi...
Xin-Hong Xie, Piyush Nagasubramaniam, Neeraj Karamchandani et al.· 0 citations
Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating scale. We show that this endpoint is not identified on the scale that rep...
Shu-Yi Fan, Boyuan Deng, Mengyu Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.