Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models
A five-stage semi-automatic framework for constructing complex graph reasoning benchmarks that serves as a challenging and diagnostic benchmark for graph reasoning and provides empirical guidance for future enhancement methods is proposed.