Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences the benchmark's measur...
Neeraj Karamchandani, Piyush Nagasubramaniam, Xin-Hong Xie et al.· 0 citations
This work presents a systematic study of honeypot-aware budget allocation for LLM attack agents and shows that with the proposed detector-guided policy, LLM agent attackers can effectively allocate budget to compromise hosts in a host pool, highlighting the importance of dynamically allocating budget in a controlled mi...
Xin-Hong Xie, Piyush Nagasubramaniam, Neeraj Karamchandani et al.· 0 citations
The results demonstrate that a clean teacher alone is not a sufficient safeguard: poisoned distillation data can produce a strongly backdoored student while maintaining competitive performance on clean images.
This work proposes Text-to-Unlearn, a novel framework that selectively unlearns concepts from pre-trained GANs using only text prompts, enabling feature and identity unlearning, as well as fine-grained tasks such as expression and multi-attribute removal in models trained on human faces.
Piyush Nagasubramaniam, Neeraj Karamchandani, Chen Wu et al.· Journal of Cybersecurity and...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.