Skip to content

SAGE: Symbolic Action-Gating and Editing for LLM Task Planners

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

SAGE (Symbolic Action-Gating and Editing), a single-LLM planner built from two lightweight mechanisms: a domain-agnostic symbolic gate that blocks precondition-violating actions with typed reasons as a runtime safety monitor, and a local edit that regenerates only the failed sub-goal's suffix, keeping completed and untouched work intact.

Abstract

Large language models (LLMs) are now the default cognitive core of embodied household agents, yet the plans they emit are rarely checked against a grounded model of the environment before execution, and the task-success they report is often measured on benchmarks so saturated that no method can be separated from another. We present SAGE (Symbolic Action-Gating and Editing), a single-LLM planner built from two lightweight mechanisms: a domain-agnostic symbolic gate (~250 lines of Python, zero tokens, $O(|\pi|)$) that blocks precondition-violating actions with typed reasons as a runtime safety monitor, and a local edit that regenerates only the failed sub-goal's suffix, keeping completed and untouched work intact; a hybrid seed+live memory store supports cold-start coverage. We evaluate under a leak-free protocol (leave-one-out retrieval) over five open-weight models and a 75-task AI2-THOR benchmark. On the standard benchmark goal-completeness saturates (52% of instances trivially solved) and SAGE ties strong hierarchical baselines. On a harder, method-agnostic multi-goal composition, SAGE's completeness lead re-emerges large (+0.06 to +0.23 across four models). Under injected mid-execution failures, SAGE recovers as reliably as whole-plan replanners at 2.4-3.3x fewer LLM calls. As a verify-before-execute gate, the symbolic monitor blocks unsafe actions before actuation and raises simulator-reported step-success for every planner tested (up to +0.11), a signal the verifier never sees (non-circular). Because the gate calls no model (0.008 ms/plan), it is a safety layer that runs essentially free on the edge: SAGE planning reproduces its quality on a Jetson AGX Orin, where small-model verification helps most. We release the benchmark, the leak-free protocol, the recovery and safety-gate harnesses, and a verifier-portability study (auto-induced on ALFWorld, 0.89 held-out).

View source

Similar papers

Preprint Sep 2026

Safe Task Planning with Long-Term Graph Memory for Embodied Agents

Large language models (LLMs) and vision-language models (VLMs) have significantly advanced zero-shot task planning for embodied agents. However, most LLM- and VLM-driven methods struggle to generate safe high-level actions due to a lack of physical risk awareness, particularly under partial observability, where hazards...

Si-Yuan Li, Tai-Yan Lang, Ao Yan et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Meta-Ctrl: Guaranteed Plan Generation by Decoupling Syntactic and Semantic Constraints

Meta-Ctrl is proposed, a constrained-decoding framework that guarantees the encoded constraints while preserving the base LM's plan quality, and is demonstrated on a real tabletop robot, where every generated plan satisfies its preconditions and goals by construction.

Gwen Yidou-Weng, Edward Sun, Tian-Yi Ma et al. · 0 citations
#artificial intelligence Preprint Aug 2026

CEDAR: Automata as Verifiable Interfaces for Language-Guided Embodied Action

CEDAR is presented, a counterexample-guided framework that grounds instructions as regular languages over environment event traces and represents both skills and specifications as deterministic finite automata, suggesting that regular languages offer a practical verification layer between natural-language instructions...

Le Chen, Alvaro Velasquez, Ashutosh Trivedi · 0 citations
#artificial intelligence Preprint Sep 2026

Symbolic Temporal Supervision of LLM Agents Using Contracts

ContrAgent is presented, a contract-based framework for symbolic temporal supervision of LLM agents that matches state-of-the-art LLM-judge and rule-based guardrail baselines while producing deterministic, reproducible verdicts and, in the online mode, orders-of-magnitude lower per-call latency.

Yi-Feng Xiao, Pierluigi Nuzzo · 0 citations
Preprint Aug 2026

Joint Optimization of Tool Creation and Use for Large Language Model Agents

A reinforcement learning framework that jointly trains tool creation and tool use inside a single policy, with three separate reward axes that catch schema, code, and outcome failures independently, so each failure mode contributes its own gradient.

Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen et al. · 2 citations
#artificial intelligence Preprint Sep 2026

How Strongly Should Task State Influence an LLM Agent?

Across three models, two reasoning regimes, and two domains, four findings hold without per-turn reasoning: displaying accurate state is unreliable, an unverified ledger the agent writes itself beats an accurate checklist it is shown, directives help in proportion to the model's obedience, and enforcement needs no obed...

C. Zhang, Wonbin Kweon, Jiawei Han · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.