Preprint
Jul 2026
AllocBench: Measuring Online Tool Allocation Capability in LLM Agents
A paired benchmark that tests whether LLM agents exhibit conscious allocation behavior under a fixed budget in two contexts: an abstract text-based formulation and a code-construction task finds that every frontier model testedacts near-optimally in the abstract framing but fails to transfer this ability to script-writing.
Daniel Wang, Andrew Xu
· 1 citation