Preprint
Aug 2026
Invocation-Level Reliability of Tool-Using Agents
This work measures a correct-invocation rate that separates the two, under both a clean teacher-forced context and the model's own free-running context, on five open-weight models over contamination-free multi-step tasks.
Afiya Noorain, Subhranshu Mohanty, Amritesh Banerjee et al.
· 0 citations