Skip to content

GTA: Graph Theory Agent and Benchmark for Algorithmic Graph Reasoning with LLMs

Sep 2026 · 2 citations · ⚡ 1 influential · 139 references
Computer Science

TL;DR

The Graph Theory Agent (GTA), which pairs a preference-trained representation selector with plan-and-decompose scaffolding around a frozen executor LLM, is proposed, which lifts Phi-4 from 53.5% to 69.1% on the benchmark's easy split and from 33.0% to 41.5% on its hard split.

Abstract

Large Language Models (LLMs) are increasingly asked to reason over structured data such as graphs, yet how reliably they can carry out multi-step graph algorithms in language remains unclear. Existing evaluations tend to use simple tasks on small graphs, to score code generation rather than reasoning over the graph itself, or to fix a single input format. We introduce Graph Theory Bench (GT Bench), a benchmark covering 24 classical graph problems in 44 task-structure settings, with over 100,000 examples across four representations: natural language, structured language, adjacency list, and adjacency matrix. Evaluating eight LLMs on GT Bench shows that accuracy is strongly tied to the input representation, that the best representation shifts with graph density, size, and topology as well as with the model, and that this sensitivity persists, attenuated, in the strongest reasoning models. Building on these observations, we propose the Graph Theory Agent (GTA), which pairs a preference-trained representation selector with plan-and-decompose scaffolding around a frozen executor LLM. GTA lifts Phi-4 from 53.5% to 69.1% on the benchmark's easy split and from 33.0% to 41.5% on its hard split, outperforming eight prompting and agent baselines, and transfers without retraining to GraCoRe and NLGraph. Code for benchmark generation and evaluation: https://github.com/xzx34/GTA. The project homepage is available at https://xzx34.github.io/gta/.

View source

Similar papers

Preprint Aug 2026

GABench: A Comprehensive Benchmark for Evaluating LLM Agents on Graph Analysis Tasks

GABench is introduced, a comprehensive benchmark for agentic graph analysis that covers four graph analysis task categories: graph retrieval, graph theory, graph machine learning, and graph open-ended question answering and provides practical insights into the development and evaluation of LLM agents for graph analysis...

Jiarui Tan, Zhong-Jian Zhang, Yabo Guo et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Multi-Agent Agentic Graph Learning via Structural Signatures

A multi-agent agentic graph learning framework that partitions the graph into communities and assigns an independent agent to each community for region-specific specialization and shows that MAAGL outperforms SOTA AGL methods.

Liang Qu, Jian-Xin Li, Hua Wang · 1 citation
#artificial intelligence Preprint Sep 2026

Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

The Procedural Graph is introduced: just as a knowledge graph organizes factual knowledge into (entity, relation, entity) triplets for what-is questions, a Procedural Graph organizes procedural knowledge into (procedure, relation, procedure) triplets for what-to-do questions.

Yu-Xing Lu, Yi-Cheng Chen, Shan-Chan Wu et al. · 0 citations
Preprint Aug 2026

Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models

A five-stage semi-automatic framework for constructing complex graph reasoning benchmarks that serves as a challenging and diagnostic benchmark for graph reasoning and provides empirical guidance for future enhancement methods is proposed.

Fali Wang, Ali Al-Lawati, Iliyas Bektas et al. · 0 citations
#machine learning Preprint Sep 2026

GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills

GraphSkillEvo is introduced, a population-based evolutionary optimization framework with mutation and crossover operators for graph-structured skills that enables broader and more comprehensive exploration of the structured skill space than purely LLM-based iterative self-refinement.

Rui Sun, Zhi Zheng, Zhen-Kun Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LADDER: Graph-Guided Diffusion Language Models for Efficient Multi-Hop Reasoning

LADDER is a novel framework that bridges diffusion language modeling with GraphRAG through graph-guided parallel decoding through graph-guided parallel decoding, and proposes an event-driven self-clocking retrieval, inspired by the key insight that 88% of target entities emerge early in the partially denoised state.

Sen-Lei Zhang, Lin-Hao Luo, Qian-Wen Zhang et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.