Network simulation plays a critical role in improving the efficiency of large-scale AI clusters for design validation, parameter tuning, and protocol development. However, high-fidelity network simulation becomes prohibitively slow at scale, especially when running large batches of experiments on topologies with tens or hundreds of thousands of accelerators. We observe that a key bottleneck comes from the control plane. Existing network simulators typically compute routes and install forwarding tables at initialization, which can consume hundreds of GB of memory before packet-event execution begins and limit overall simulation throughput. In this paper, we present Nüwa, which views routing as a compilation problem, it leverages the hierarchical and symmetric structure common in AI fabrics and compiles a declarative topology description together with routing policies into compact forwarding artifacts that are fast to generate and efficient to look up. Evaluations show that Nüwa can reduce simulation initialization time from hours to only 25 seconds for a 65,536-GPU cluster. For end-to-end simulation time, Nüwa takes only 20% of that required by existing approaches in a 40K+ GPU cluster, and Nüwa can scale to a 221,184-GPU cluster.
Wenkai Li, Ran Shu, Peng Zhang et al.· Proceedings of the ACM SIGCO...· 0 citations
Quanta is presented, a redundancy-aware simulation framework that eliminates repeated execution in packet-level simulation and reduces the dependence of simulation cost on physical network scale.
Jiajun Luan, Hao Li, Yihan Dang et al.· Asia-Pacific Workshop on Net...· 0 citations