Skip to content
Book Open access

Investigating Reasoning in Large Language Models with Counterfactual Knowledge Graphs

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · 1 citation · 35 references

TL;DR

This work delineates LLM reasoning boundaries and presents a new paradigm for fine-grained capability assessment, which suggests that genuine reasoning is demonstrated only when a model follows logical rules despite conflicting prior knowledge.

Abstract

Despite the success of Large Language Models (LLMs) on reasoning benchmarks, it remains unclear whether their performance stems from genuine logical deduction or the memorization of training patterns. Existing benchmarks often fail to disentangle reasoning from prior knowledge, as tasks grounded in real-world facts allow models to take ''knowledge shortcuts''. In this paper, we propose a novel diagnostic benchmark to decouple knowledge memorization from logical reasoning. Built on the DBpedia KG, our framework constructs multi-hop reasoning chains (from Q1 to Q5) across three task dimensions: Factural Questions (FQ), Counterfactual Questions (CQ) with logically consistent but counterfactual conclusions, and Questions with Similar-Entity Options (SO) to evaluate the dependence on prior knowledge. Questions without Context serve only as an intermediate form: they contain solely queries with no triples or options, so LLMs cannot answer them directly. Our core hypothesis is that genuine reasoning is demonstrated only when a model follows logical rules despite conflicting prior knowledge. Evaluating seven state-of-the-art LLMs (8B to ultra-large) reveals strong prior knowledge dependence, with performance degrading sharply on counterfactual tasks as reasoning depth grows. This work delineates LLM reasoning boundaries and presents a new paradigm for fine-grained capability assessment.

Read PDF

Similar papers

Preprint Aug 2026

Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge

TKFQA, a factuality consistency benchmark comprising 10,130 question-answering (QA) pairs grounded in tables, texts, and knowledge graphs, is introduced and ORLF, an LLM-agnostic training framework that models cross-context topological relations through knowledge-specific latent vectors is proposed.

Shibo Chu, Yu-Ze Liu, Tie-Hua Zhang et al. · 0 citations
Preprint Aug 2026

NeSy-RAG: Neuro-Symbolic RAG for Explainable Question Answering

This work proposes NeSy-RAG, a modular neuro-symbolic RAG framework that synthesizes attributable Prolog modules from retrieved text chunks that introduces a symbolic knowledge-gap detection mechanism that identifies missing user facts whose truth values affect the query outcome and automatically triggers follow-up int...

Jonas Gann, Michael Gertz · 1 citation
#large language models Open access Sep 2026

Emergent Reasoning in Large Language Models: A Systematic Evaluation Across Task Complexity

Large language models (LLMs) are said to exhibit “emergent” reasoning capabilities — ones that are virtually nonexistent in smaller models but suddenly emerge as soon as the model size surpasses a critical point. This claim has been at the heart of discussions on the capability forecasting, safety planning and evaluati...

N. Kuotsu · 0 citations
#artificial intelligence Preprint Sep 2026

The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA

Small language models (SLMs) are increasingly paired with knowledge graphs (KGs), yet end-to-end KG question answering conflates graph access, search, navigation, reasoning, and answer generation. This coupling makes it difficult both to determine whether an SLM can faithfully execute the reasoning path implied by a qu...

Eduin E. Hernandez, Sergio A. Diaz, Luis F. Garcia et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond Surface Forms: Symbolic Edits as a Test for Logical Reasoning with LLMs

Logical reasoning with large language models (LLMs) is a critical capability, as it reflects a system's ability to correctly deduce hypotheses from a given context using faithful deductive processes. However, LLM reasoning has often been shown to be sensitive to small surface-level variations in problem formulation, ra...

Ramya Keerthy Thatikonda, W. Buntine, Ehsan Shareghi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.