Preprint
Aug 2026
NetConfArena: An Executable Benchmark for LLM Agents in Closed-Loop Network Configuration
NetConfArena is presented, an executable benchmark for evaluating LLM agents in closed-loop network configuration, and its findings suggest two future directions: using validated trajectories as supervision signals to improve foundation models, and designing harness mechanisms that make agent execution more reliable and accountable.
Chang Liu, Xiaohui Xie, Xinyi Chen et al.
· 0 citations