Preprint
Jul 2026
ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management
ARBIGRAPH is introduced, a benchmark generator for evaluating whether tool-assisted language agents can retain, update, compose, and discard task-relevant context across extended reasoning workflows, and shows that ARBIGRAPH exposes failures that are not visible from single-task evaluation alone.
Pavel Golikov, E. Opryshko, Gennady Pekhimenko et al.
· 0 citations