Shortcutting the Fix: Identifying and Categorizing Agentic Exploits in Software Engineering Benchmarks
While autonomous software engineering (SWE) agents achieve high benchmark resolution rates, these scores can mask exploitative behaviors---such as leveraging local Git histories, accessing upstream repositories, or recalling memorized solutions---rather than demonstrating genuine problem solving. We systematize and aud...