Traditional application architectures assume behavioral logic authored in advance, leaving reachable behavior largely bounded by explicit code and workflow rules. Large language model (LLM)-based agents challenge this assumption by enabling runtime reasoning and autonomous action to become part of application behavior, allowing applications to address situations not enumerated at design time. Extending this pattern, this paper identifies a new application paradigm in which LLM-based agents serve as central reasoning and action components responsible for the application’s core logic and, where permitted, for adapting the application graph itself at runtime. We call these agent-native applications. While such applications significantly expand their possible behavioral space beyond explicit code and workflow rules, they also face a major control problem in which useful agentic reasoning should be preserved while application behavior should remain within a permissible space. We therefore propose an architectural model that represents the application as a portable graph of agents, tools, data sources, and human-in-the-loop (HITL) checkpoints, and encodes the application’s permitted behavior as a behavioral envelope within a declarative application specification. At runtime, an application orchestrator serves as the control plane that coordinates tasks and governs how the graph and its permissions evolve, while an agent mesh serves as the data plane that mediates policy-relevant interactions and produces audit events. We then discuss the trust layer that makes agent-native applications governable and the supporting foundations required for practical operation. Two use cases illustrate the architecture, while Agent-Native Runtime (ANR) demonstrates selected core mechanisms in an executable prototype.
LLM agents increasingly operate as execution systems that invoke tools, modify local state, use persistent memory, and interact with external protocols. These capabilities make agents useful, but they also introduce risks related to over-privileged actions, weak auditability, prompt injection, tool poisoning, and uncontrolled side effects. This paper presents Agentao, a governed local-first runtime for tool-using LLM agents. Agentao separates model-generated action proposals from host-authorized execution through a layered architecture consisting of host-facing surfaces, a host contract, a runtime core, a permission-mediated tool system, and supporting subsystems for memory, replay, plugins, skills, sub-agents, and protocol integration. We describe the motivation, threat model, design goals, governance model, execution pipeline, and structured event interface of the system. Agentao does not provide formal safety guarantees; rather, it demonstrates how permissions, state, protocol boundaries, and execution traces can be made explicit runtime abstractions for building agents that are more governable, inspectable, and suitable for host-controlled local environments. The code is publicly available at https://github.com/jin-bo/agentao.
AgentFlow, a flow-centric policy language and runtime enforcement model for specifying where data may travel in agent systems, is presented and results are preliminary and scoped to the modeled policy-visible agent behaviors and evaluated benchmarks.
A code agent-agnostic agentic scaffold for automated code performance optimization that enables the system to autonomously execute the entire pipeline: project-level runtime analysis, hotspot identification and benchmark extraction, Abstract Syntax Tree (AST)-precise code localization, candidate patch generation, functional verification, performance measurement, and version rollback.
Agent Gym is introduced, a modular, domain-agnostic framework that wraps any existing LLM-based agent in a continuous evaluation-and-evolution loop and introduces the Spec-to-Note Gap, an autoencoder-inspired view of agentic system transparency.
Pouya Ghiasnezhad Omran, Michael Zimmermann, Duncan Cambridge et al.· 0 citations
A design pattern is introduced that lets a coding agent work inside the environment directly, and marimo’s rules accumulate the agent’s exploration into a runnable, reproducible Python program.