Study of an end-to-end NL-to-PDDL pipeline that combines LLM generation, checks in terms of PDDL parsing, planning and validation, a domain-conformance checker, an LLM critic, and iterative repair shows that structured repair can be useful, and that PDDL~2.1 remains challenging for reference reconstruction, even when operational success improves.
Abstract
Large Language Models (LLMs) have shown promise for translating Natural Language (NL) planning descriptions into PDDL problem instances. However, standard evaluation criteria such as syntactic validity or planner success can substantially overestimate faithfulness to the described task: a generated problem may be parseable and solvable while misrepresenting the intended initial state, goal, object structure, or optimization target. This paper studies an end-to-end NL-to-PDDL pipeline that combines LLM generation, checks in terms of PDDL parsing, planning and validation, a domain-conformance checker, an LLM critic, and iterative repair. Fine-grained repair feedback is constructed from the domain description, the generated problem, the natural language problem description, and operational diagnostics. Reference-based comparisons against curated benchmark PDDL problem descriptions are used for post-hoc benchmark analysis, and these offline checks include renaming-invariant structural matching and semantic equivalence, where domain support is available. Across Planetarium, AutoPlanBench, and curated PDDL~2.1 problems, results show that operational success and benchmark-reference reconstruction can diverge substantially. Results also show that structured repair can be useful, and that PDDL~2.1 remains challenging for reference reconstruction, even when operational success improves.
GenCTL is proposed, a prompt-based framework for effective and model-aware NL2CTL translation without task-specific fine-tuning that improves the reliability and practical checkability of LLM-generated CTL specifications.
Ran Tao· Poster Volume 0007 The 2026...· 0 citations
PDDLCoder is presented, an agentic framework for PDDL generation from natural language that iteratively generates, analyzes, and refines planning specifications and demonstrates the effectiveness of agentic PDDL generation for planning and establishes a reproducible benchmark for future research on LLM-assisted symboli...
Veit Laule, Jiangtao Shuai, Manfred Hauswirth et al.· 0 citations
LLMSuite is proposed, a hybrid test generation framework that integrates self-refinement prompting with class-level LLM reasoning into the search-based testing process and complements manually written test suites by exercising domain-specific behaviors that are often left untested.
Amirhossein Deljouyi, Annibale Panichella, Andy Zaidman· 0 citations
This paper empirically evaluates whether current Large Language Models can reliably translate natural language testing goals, written in informal language by video game testers, into well-formed PDDL targets suitable for classical planning.
Tomás Balyo, L. Chrpa, G. M. Youngblood· 0 citations
Pseudo2Code, a benchmark designed to systematically evaluate the impact of structured pseudocode on code generation quality and algorithmic faithfulness, is introduced and the Pseudo2Code Agentic Framework is proposed, a multi-stage pipeline that leverages pseudocode as an explicit intermediate reasoning representation...
Shadikur Rahman, Umme Ayman Koana, S. Danish· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.