Skip to content

Grounded Evaluation and Repair for NL-to-PDDL Problem Generation

Sep 2026 · 0 citations · 18 references
Computer Science

TL;DR

Study of an end-to-end NL-to-PDDL pipeline that combines LLM generation, checks in terms of PDDL parsing, planning and validation, a domain-conformance checker, an LLM critic, and iterative repair shows that structured repair can be useful, and that PDDL~2.1 remains challenging for reference reconstruction, even when operational success improves.

Abstract

Large Language Models (LLMs) have shown promise for translating Natural Language (NL) planning descriptions into PDDL problem instances. However, standard evaluation criteria such as syntactic validity or planner success can substantially overestimate faithfulness to the described task: a generated problem may be parseable and solvable while misrepresenting the intended initial state, goal, object structure, or optimization target. This paper studies an end-to-end NL-to-PDDL pipeline that combines LLM generation, checks in terms of PDDL parsing, planning and validation, a domain-conformance checker, an LLM critic, and iterative repair. Fine-grained repair feedback is constructed from the domain description, the generated problem, the natural language problem description, and operational diagnostics. Reference-based comparisons against curated benchmark PDDL problem descriptions are used for post-hoc benchmark analysis, and these offline checks include renaming-invariant structural matching and semantic equivalence, where domain support is available. Across Planetarium, AutoPlanBench, and curated PDDL~2.1 problems, results show that operational success and benchmark-reference reconstruction can diverge substantially. Results also show that structured repair can be useful, and that PDDL~2.1 remains challenging for reference reconstruction, even when operational success improves.

View source

Similar papers

Preprint Aug 2026

PDDLCoder: Agentic PDDL Generation for LLM-Assisted Symbolic Planning

PDDLCoder is presented, an agentic framework for PDDL generation from natural language that iteratively generates, analyzes, and refines planning specifications and demonstrates the effectiveness of agentic PDDL generation for planning and establishes a reproducible benchmark for future research on LLM-assisted symboli...

Veit Laule, Jiangtao Shuai, Manfred Hauswirth et al. · 0 citations
#software testing Preprint Sep 2026

Enhancing Automated Unit Test Generation for NLP Libraries Using Large Language Models

LLMSuite is proposed, a hybrid test generation framework that integrates self-refinement prompting with class-level LLM reasoning into the search-based testing process and complements manually written test suites by exercising domain-specific behaviors that are often left untested.

Amirhossein Deljouyi, Annibale Panichella, Andy Zaidman · 0 citations
#artificial intelligence Preprint Sep 2026

Which LLM is Best for Translating Natural Language Goals to PDDL

This paper empirically evaluates whether current Large Language Models can reliably translate natural language testing goals, written in informal language by video game testers, into well-formed PDDL targets suitable for classical planning.

Tomás Balyo, L. Chrpa, G. M. Youngblood · 0 citations
Preprint Aug 2026

Pseudo2CodeQA: A Benchmark for LLM-Based Structured Algorithmic Reasoning in Code Generation

Pseudo2Code, a benchmark designed to systematically evaluate the impact of structured pseudocode on code generation quality and algorithmic faithfulness, is introduced and the Pseudo2Code Agentic Framework is proposed, a multi-stage pipeline that leverages pseudocode as an explicit intermediate reasoning representation...

Shadikur Rahman, Umme Ayman Koana, S. Danish · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.