Preprint
Jul 2026
Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI
What the results imply for how such orchestrators should be evaluated are discussed, and the study uses a single benchmark, the fixes were tuned on the same challenges they were evaluated on, and the client effect is demonstrated for one model only, so its generality to other models remains a hypothesis.
Romain Gerard, Assmaa Zeghaider, Yan Guo
· 0 citations