Skip to content

TopoCrafter: Toward Highly Reliable Optical-Circuit-Switched HPC/AI Interconnects via Dual-Agent DRL

2026 · IEEE Transactions on Cognitive Communications and Networking · Vol 12, pp. 11277-11292 · 0 citations · 49 references

Abstract

Optical-circuit-switched interconnects have become one of the core components for AI training due to their flexible topology reconfiguration. In contrast to the applications carried by traditional data center networks, large-scale language model training is highly sensitive to network failures, where frequent disruptions will cause gradient synchronization delays, leading to training interruptions and wasted computational resources. Existing schemes are primarily focused on specific communication patterns, without considering the fault probability distribution. As a result, unreliable links remain on critical paths. Furthermore, passive fault response mechanisms lead to inefficient topology reconfigurations, preventing network protocol convergence and making it difficult to meet the stringent stability requirements of large-scale model training. To address reliability challenges in optical-circuit-switched interconnect, we propose TopoCrafter, which leverages dual-agent deep reinforcement learning to proactively mitigate network failures. The “Topo-Agent” estimates link failure probabilities to determine reconfiguration timing and then employs a lightweight heuristic algorithm to create failure-avoidant topology that matched to traffic pattern. Concurrently, the “Route-Agent” optimizes traffic distribution. Through their strategic interaction, the agents learn holistic policies that optimally balance network reliability and communication efficiency. To improve generalization, a progressive training approach is employed, allowing the agents to adapt to complex failure environments while accelerating convergence. Under link failure scenarios, TopoCrafter maintains reliability, reducing end-to-end latency by up to 50% and maximum link utilization by approximately 20% compared to FatTree. In addition, progressive training algorithm ensures a performance degradation of less than 10% when adapting to new failure environments, and it maintains stable high performance as the network scales.

View source