Skip to content

Which LLM is Best for Translating Natural Language Goals to PDDL

Sep 2026 · 0 citations · 20 references
Computer Science

TL;DR

This paper empirically evaluates whether current Large Language Models can reliably translate natural language testing goals, written in informal language by video game testers, into well-formed PDDL targets suitable for classical planning.

Abstract

Bridging the gap between human intent and machine execution remains a challenge in automated planning, where expressing goals in formal languages like PDDL restricts accessibility to non-experts. This paper empirically evaluates whether current Large Language Models (LLMs) can reliably translate natural language testing goals, written in informal language by video game testers, into well-formed PDDL targets suitable for classical planning. We present a carefully designed prompt template, integrating insights from iterative experimentation, aimed at maximizing both accuracy and response coherence from multiple state-of-the-art LLMs. Six contemporary models are systematically assessed on correctness, speed, and error tendencies using real-world, domain-specific benchmarks. All models demonstrate high correctness, exceeding 92\%, with Gemini 2.5 Flash achieving the highest accuracy at 96\% and the lowest incidence of false positives, while GPT-4.1 leads in response speed. Despite these advances, critical distinctions exist in model performance, and occasional failures arise from language ambiguity and limitations in domain representation. Our analysis underscores both the significant progress and ongoing gaps in enabling LLMs to act as robust bridges between natural language objectives and automated planning pipelines.

View source

Similar papers

Preprint Aug 2026

LLM-Only PDDL Domain Repair with Open-Weight Models

Current open-weight models cannot guarantee satisfaction of the test constraints required for reliable automated model repair, and this paper evaluates the ability of recent open-weight large language models to perform this repair task using an LLM-only approach.

Nader Karimi Bavandpour, Pascal Bercher · 0 citations
Open access Sep 2026

From Voice to Robot: Evaluating Large Language Models for Industrial Programming with HARPA

Recent advances in generative artificial intelligence and large language models (LLMs) have increased the feasibility of translating natural-language instructions into executable robot programs, reducing the technical barrier separating shop-floor operators from industrial robotic programming. However, current evalua...

Beatriz Consciência, Ricardo Correia, C. Wanzeller · 0 citations
#artificial intelligence Preprint Sep 2026

Grounded Evaluation and Repair for NL-to-PDDL Problem Generation

Study of an end-to-end NL-to-PDDL pipeline that combines LLM generation, checks in terms of PDDL parsing, planning and validation, a domain-conformance checker, an LLM critic, and iterative repair shows that structured repair can be useful, and that PDDL~2.1 remains challenging for reference reconstruction, even when o...

J. Rosa, Pedro Santos, Valdemar Oliveira et al. · 0 citations
#artificial intelligence Preprint Sep 2026

What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

This work systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026, using staged screening and automated full-text coding to examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms.

Chao Wang · 0 citations
Open access

Evaluation and Distillation of Source Code Generation Tasks by Large Language Models

Two novel contributions are introduced: CodeEval and CodeQual, an open-source execution framework that provides researchers with a ready-to-use evaluation pipeline for evaluating and improving LLMs in software engineering contexts, encompassing both functional correctness assessment and subjective code quality evaluati...

Danny Brahman · 0 citations
#natural language process... Preprint Sep 2026

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

Recently, the rapid development of large language models (LLMs) has reshaped software engineering by enabling autonomous code agents that plan, execute, and utilize external tools iteratively to tackle complex tasks. Beyond achieving functional correctness, these agents must faithfully follow process instructions and c...

Bo-Si Wen, Cun-Xiang Wang, Jia-Yi Gui et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.