Results show the feasibility of delegating adaptive reasoning to the cloud while retaining execution control at the edge, and closed-loop repair improves application success from 0/6 to 6/6 over one-shot generation, validation gating prevents all three observed target-scale failures, and skill conditioning improves success from 4/6 to 6/6.
Abstract
Scientific applications increasingly rely on high-performance computing (HPC), yet translating a scientist's high-level goal into a correct target-scale execution remains brittle and labor-intensive. Large language model (LLM) agents promise to automate this, but two obstacles remain: granting a cloud-hosted model direct HPC access exposes credentials and execution authority, while withholding it demands continuous human supervision; and one-shot generation cannot adapt when generated artifacts fail in a site-specific HPC environment. We present \textsc{ECAS}, an \textbf{E}dge-\textbf{C}ontrolled \textbf{A}gentic \textbf{S}ystem for closed-loop execution of scientific computing campaigns with limited human intervention. \textsc{ECAS} separates \emph{reasoning}, \emph{control}, and \emph{execution}: a cloud-hosted LLM proposes plans, artifacts, and repairs; a user-controlled edge agent retains credentials, workflow state, and execution authority while enforcing policy and resource constraints; and the HPC system computes. Its core mechanism is \emph{validation-gated execution}: generated artifacts pass static checks and small-scale validation, failures trigger repairs from sanitized execution feedback, and target-scale execution is permitted only after validation and policy checks pass. \textsc{ECAS} also draws on an edge-resident library of expert-distilled, site-specific skills that is never disclosed to the cloud. In preliminary experiments with three scientific applications on two production ALCF systems under six injected fault types, closed-loop repair improves application success from 0/6 to 6/6 over one-shot generation, validation gating prevents all three observed target-scale failures, and skill conditioning improves success from 4/6 to 6/6. These results show the feasibility of delegating adaptive reasoning to the cloud while retaining execution control at the edge.
Large language model agents can automate data science workflows, but cloud-centric deployment exposes sensitive context and edge-only deployment limits analytical capability. We present FinDS-Agent, a cloud–edge framework that keeps raw records and program execution at the trusted edge while providing a policy-screened...
Xiao-Zheng Du, Rui-Jun Deng, Cheng Wang et al.· Future Internet· 0 citations
The Model Context Protocol (MCP) provides a common interface through which AI applications discover and use external resources and tools. It allows language-model agents to ground their reasoning in current system state and interact with heterogeneous services. In medical environments, however, exposing device state an...
Apodex 1.1 reaches the leading performance band despite using a substantially smaller model than many frontier systems, and the 35B-parameter Apodex 1.1 Mini further retains strong working capability in a locally deployable form.
Scientific workflow management (WMSs) systems automate execution, yet orchestrate using fixed, hand-tuned rules. LLM agents promise more autonomous orchestration, but it remains unclear where to introduce agentic reasoning, how to bound its risk, and when it actually helps. We present Avatar, an actor-based architectur...
Suman Raj, H. Nguyen, Haochen Pan et al.· 0 citations
AI agents are emerging as a practical way to run multi-step scientific workflows that interleave reasoning, tool use, and verification. Scaling such agentic science remains difficult because workflows are hard to observe and reproduce, many scientific tools and laboratory systems are not agent-ready, and execution trac...
Lin-Feng Zhang, Si-Heng Chen, Yu-Zhu Cai et al.· AI Plus· 0 citations
The proliferation of large language model (LLM)-based autonomous agents has created a new class of distributed system: the multi-agent LLM network. While significant research focuses on the intelligence of individual agents, comparatively little work addresses the software architectural concerns that govern how fleets...
Ketankumar Savajiyani· 2026 International Conferenc...· 0 citations