Aug 2026· IEEE Transactions on Pattern Analysis and Machine Intelligence· Vol PP, pp. 1-20· 0 citations
Medicine
TL;DR
This work presents an exposition of the OSQ problem by summarizing its various formulations in the current literature and categorizing existing solutions into three different types, and summarizes the empirical methods proposed by existing works to verify the efficiency of OSQ mitigation approaches.
Abstract
Graph-based message-passing neural networks (MPNNs) have achieved remarkable success in both node and graph-level tasks. However, several identified problems, including over-smoothing (OSM), limited expressive power, and oversquashing (OSQ), still restrain the performance of MPNNs. In particular, the latest identified problem, OSQ, reveals that MPNNs generally fail to maintain their learning accuracy with tasks that require long-range dependencies between node pairs. In this work, we present an exposition of the OSQ problem by summarizing its various formulations in the current literature and categorizing existing solutions into three different types. In addition, we also discuss the alignment between OSQ and expressive power plus the trade-off between OSQ and OSM. Furthermore, we summarize the empirical methods proposed by existing works to verify the efficiency of OSQ mitigation approaches, together with illustrations of their computational complexities. Lastly, we identify some open questions that are of interest for further exploration of the OSQ problem.
The design of transformer-based Large Language Models (LLMs) is being radically changed through new architectures that are able to overcome scalability limitations of previous designs, including Mixture-of-Experts (MoE), Multi-Head Latent Attention (MLA), and Multi-Token Prediction (MTP). As an open-weighted model rele...
Yassine Zouhdi, B. Hdioud· EPJ Web of Conferences· 0 citations
Achieving strong performance with graph neural networks (GNNs) typically requires training and hyperparameter tuning for each dataset, incurring repeated costs and effort. Graph in-context learning (ICL) avoids this by using a single pretrained model to predict unknown node labels directly from labeled context nodes. E...
Dooho Lee, Jin-Mo Lee, Minho Jeong et al.· 0 citations
Graph distribution shifts between training and test graphs pose severe challenges to the generalization of graph neural networks (GNNs). In real-world deployment, application environments are continuously evolving, while retraining or redesigning GNNs is often costly and impractical. In light of this, test-time adaptat...
Jiayi Chen, Xin Zheng, Bo Li et al.· Proceedings of the Thirty-Fi...· 3 citations
A differentiable Sandpile Stabilization Layer (SSL) and congestion-aware objectives designed to redistribute excess load and manage stabilization costs are proposed and Experiments on long-range benchmarks show that targeting sandpile-identified bottlenecks mitigates representation collapse and improves over standard b...
Yang Shi, Li-Xian Chen, Jingchao Wang et al.· Proceedings of the 32nd ACM...· 0 citations
This study introduces a novel memory-augmented self-learning framework that extracts and provides diverse learning sources for adaptive knowledge distillation from the student model itself, resulting in a 2.5-6% increase in accuracy across various benchmark datasets compared to current GNN training and self-distillatio...
Saurabh Sharma, Souvik Chowdhury, Joydeep Chandra· Data mining and knowledge di...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.