Assurance-Aware Semantic Scheduling (AAS) combines integer-program selection, dispatch-aware temporal scheduling, bounded diagnostic expansion, and receipt-aware repair to formulate evidence acquisition as joint witness selection and scheduling under quorum, diversity, freshness, deadline, and resource constraints.
Abstract
Tool-using agents can initiate consequential infrastructure changes, yet evidence required for admission may expire while other checks run or depend on a shared fault domain. We formulate evidence acquisition as joint witness selection and scheduling under quorum, diversity, freshness, deadline, and resource constraints. Assurance-Aware Semantic Scheduling (AAS) combines integer-program selection, dispatch-aware temporal scheduling, bounded diagnostic expansion, and receipt-aware repair. Formal results state the assumptions needed for dispatch-time freshness and finite diagnostic expansion. In three generated infrastructure workloads, AAS produces 1,075/1,200 valid candidates versus 647/1,200 for constraint-aware forward scheduling; stale candidates fall from 440 to 12. Paired sensitivity studies reuse the same instances and operation latency draws across parameter settings. A corrected timeout intervention finds 18/20 admissions with repair or full resynthesis versus 0/20 for a static plan, with lower committed cost when receipts are reused. On 20 constructed cases requiring a certified decomposition cut, refinement recovers an oracle-matching feasible plan every time. These are controlled simulation results; the bounded oracle shares a temporal search component, and transfer to deployed systems remains untested.
ClawProBench is presented, a trace-aware benchmark for runtime-native agent evaluation instantiated on OpenClaw, a live agent runtime with workspace tools and native surfaces for browsing, memory, messaging, scheduling, skills, and subagents.
A tool-using model can follow a malicious instruction even when its credentials are valid. We study whether task-scoped authorization contains the resulting tool execution. Our paired-replay testbed samples a model request once and submits the same action, resource, and arguments to broad bearer, scoped JWT, sender-con...
Conventional reasoning protocols present a fixed, preselected task, so they cannot test whether an agent propagates relevant updates, preserves unaffected work, or reconstructs a historical task binding. We therefore study \emph{dynamic task routing}, in which an event stream revises task bindings and a system must sel...
Yang-Hao Luo, Selim-Antoine Lali, Jeremy Moebel et al.· 0 citations
Evidence-Carrying Termination (ECT): an agent may return COMPLETE only when a typed certificate binds every required answer claim to valid, in-scope trace evidence and a deterministic replay reconstructs the claimed value.
Agents based on large language models (LLMs) can access heterogeneous devices through tools and APIs, but reliable execution must account for unmet effects, uncertain outcomes, and changing prerequisites. A command may be acknowledged without producing its intended effect, while missing feedback may obscure an action t...
Xue-Chun Li, Jia-Xin Liang, Jie Li et al.· 0 citations
This work formalizes the admission calculus and the assumptions connecting it to mediated execution, and establishes tested implementation behaviors and local costs, not production failure rates or comparisons of language-model capability.
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.