Current open-weight models cannot guarantee satisfaction of the test constraints required for reliable automated model repair, and this paper evaluates the ability of recent open-weight large language models to perform this repair task using an LLM-only approach.
Abstract
AI planning is concerned with finding a sequence of actions that achieves a specified goal. It relies on explicit models of the world, commonly represented in the Planning Domain Definition Language (PDDL). An active line of research investigates how errors in such models can be detected and repaired. For example, users may provide positive test plans that are solutions, and negative test plans that fail during execution. Automated repair methods then modify the PDDL model to satisfy these constraints. In this paper, we evaluate the ability of recent open-weight large language models to perform this repair task using an LLM-only approach. Our experiments show that the symbolic baseline achieves an $F_1$ score of $.49$, while the best-performing LLM reaches $.87$ with high reasoning effort, an absolute improvement of $.38$. However, that setting has a mean test pass rate of only $.82$, falling to $.06$ on the Thoughtful domain; even the best setting that includes the test traces reaches only $.92$. Thus, current open-weight models cannot guarantee satisfaction of the test constraints required for reliable automated model repair.
This paper empirically evaluates whether current Large Language Models can reliably translate natural language testing goals, written in informal language by video game testers, into well-formed PDDL targets suitable for classical planning.
Tomás Balyo, L. Chrpa, G. M. Youngblood· 0 citations
This work introduces a semantic-preserving PDDL-to-Lean conversion, and uses an LLM to generate both the generalized plan and the formal proof that it solves every instance satisfying the domain constraints, and evaluates this approach on 13 commonly used benchmark domains.
Katharina Stein, Chaahat Jain, J. Hoffmann et al.· 0 citations
SAGE (Symbolic Action-Gating and Editing), a single-LLM planner built from two lightweight mechanisms: a domain-agnostic symbolic gate that blocks precondition-violating actions with typed reasons as a runtime safety monitor, and a local edit that regenerates only the failed sub-goal's suffix, keeping completed and unt...
T. Bui, Jongsul Moon, Youngouk Kim et al.· 0 citations
Empirical evaluations show that GRASP consistently establishes a new state-of-the-art frontier across diverse datasets, yielding substantial accuracy gains over direct LLM planners on Natural Plan Calendar Scheduling, ZebraLogic, and SciBench Math.
Arunabh Srivastava, M. Khojastepour, Srimat T. Chakradhar et al.· 0 citations
PDDLCoder is presented, an agentic framework for PDDL generation from natural language that iteratively generates, analyzes, and refines planning specifications and demonstrates the effectiveness of agentic PDDL generation for planning and establishes a reproducible benchmark for future research on LLM-assisted symboli...
Veit Laule, Jiangtao Shuai, Manfred Hauswirth et al.· 0 citations
This work leveraging LLMs to generate instance-generation programs, with built-in soundness guarantees through prescribed checks, and shows that these automatically generated instance generators return large numbers of sound and diverse instances efficiently.
Nicola J. Müller, Naya Rudolph, Katharina Stein et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.