SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, an...
Pujun Zheng, Zi-Xin Shang, Shufan Jiang et al.· 1 citation
AgentCompass is introduced, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents that organizes the evaluation process around three independent components, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic.
Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings.
Lei Bai, Jiaqi Cao, Chiyu Chen et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.