SASTFormer: a traffic flow prediction method based on spatial-temporal multi-head self-attention mechanism fusion
Abstract
Accurately anticipating traffic volume is essential for optimizing navigation routes, lowering fuel usage, and enhancing overall travel efficiency. However, current spatiotemporal fusion techniques often neglect how distant nodes affect local traffic conditions and fail to capture long-sequence temporal dependencies. To overcome this limitation, we propose SASTFormer, which is a method for traffic flow prediction that relies on fusing spatiotemporal multi-head self-attention. The proposed network makes use of a Transformer encoder and initially extracts temporal dependencies through a temporal self-attention mechanism, followed by a spatial self-attention component that incorporates a graph mask matrix— integrating node similarity and spatial attention—to model long-range spatial dependencies. Finally, the two modules are fused to enhance long-term prediction. Providing a theoretical foundation for urban traffic guidance, our method not only outperforms eight baseline models in overall performance and medium‑/long‑term prediction on PeMS08, but also delivers more stable predictive accuracy.