Reliable data movement is essential to the Worldwide LHC Computing Grid. This project asks whether FTS queues that share a storage endpoint contain useful information about one another’s future state. I process 52,037,899 raw queue records from January 2026 into 23,820,642 sparse queue states sampled every 2 minutes. At each prediction time, observed storage endpoints form graph nodes and directed FTS queues form temporal edges. A convolutional neural network (CNN) followed by a long short-term memory (LSTM) network first encodes the previous 20 minutes of every queue independently. One simple message-passing layer then averages incident edge embeddings at each endpoint and returns the source and destination context to the target edge. The resulting graph neural network (GNN), a parameter-matched multilayer perceptron (MLP), and a degree-preserving random graph are compared to isolate the effect of real WLCG endpoint assignment. On the final test period, real topology did not give a convincing advantage for throughput regression: random topology performed at least as well, and persistence retained the lowest mean absolute error. In contrast, the real GNN reached 0.5704 ± 0.0049 average precision for 20-minute bad-link onset, compared with 0.5212 ± 0.0031 for random topology. As a sanity check, the GNN was also compared with two simple rules based on the current bad states at the two endpoints. The stronger rule reached only 0.2627 AP. The graph advantage remained in direct and autoregressive bad-state forecasts up to 60 minutes. Matched seven-input regressions gave a target-dependent result: real topology improved success-rate MSE, while throughput showed no clear graph advantage when the inputs and selected queue windows were kept the same. This makes an explanation based only on classification being easier less likely. The results suggest that endpoint context is useful for degradation-related FTS controller quantities, but not clearly for workload-driven throughput. This is a first step towards fault-propagation modelling. The current model predicts only queues observed at the forecast origin; it does not predict future queue appearance or disappearance.
Pavel Khudov Yakovlev, Maria del Carmen Misa Moreira, Sofia Vallecorsa· Zenodo (CERN European Organi...· 0 citations
Reliable data movement is essential to the Worldwide LHC Computing Grid. This project asks whether FTS queues that share a storage endpoint contain useful information about one another’s future state. I process 52,037,899 raw queue records from January 2026 into 23,820,642 sparse queue states sampled every 2 minutes. At each prediction time, observed storage endpoints form graph nodes and directed FTS queues form temporal edges. A convolutional neural network (CNN) followed by a long short-term memory (LSTM) network first encodes the previous 20 minutes of every queue independently. One simple message-passing layer then averages incident edge embeddings at each endpoint and returns the source and destination context to the target edge. The resulting graph neural network (GNN), a parameter-matched multilayer perceptron (MLP), and a degree-preserving random graph are compared to isolate the effect of real WLCG endpoint assignment. On the final test period, real topology did not give a convincing advantage for throughput regression: random topology performed at least as well, and persistence retained the lowest mean absolute error. In contrast, the real GNN reached 0.5704 ± 0.0049 average precision for 20-minute bad-link onset, compared with 0.5212 ± 0.0031 for random topology. As a sanity check, the GNN was also compared with two simple rules based on the current bad states at the two endpoints. The stronger rule reached only 0.2627 AP. The graph advantage remained in direct and autoregressive bad-state forecasts up to 60 minutes. Matched seven-input regressions gave a target-dependent result: real topology improved success-rate MSE, while throughput showed no clear graph advantage when the inputs and selected queue windows were kept the same. This makes an explanation based only on classification being easier less likely. The results suggest that endpoint context is useful for degradation-related FTS controller quantities, but not clearly for workload-driven throughput. This is a first step towards fault-propagation modelling. The current model predicts only queues observed at the forecast origin; it does not predict future queue appearance or disappearance.
Pavel Khudov Yakovlev, Maria del Carmen Misa Moreira, Sofia Vallecorsa· Zenodo (CERN European Organi...· 0 citations