This research aims to bridge the gap by providing a comprehensive evaluation methodology for LLM-powered agents that is grounded in real-world software development practice and focuses on contamination-awareness, in thewild agentic behavior assessment, and trajectory-aware benchmarks and metrics.
The first Poisoned Chalice of LLM Evaluation Competition is organized, which frames contamination detection as a white-box membership inference task on source code and provides participants with curated datasets, target models, baseline attacks, and a final evaluation on a held-out model and dataset.
J. Katzy, Ali Al-Kaswan, R. Popescu et al.· SIGSOFT FSE Companion· 1 citation· ⚡1