Coding agents have become real users of high-performance computing (HPC) systems, yet today's HPC abstractions, interfaces, and policies remain designed for human-driven workflows. In our measurement, users running coding agents are only 19.5% of the observed population, but account for 55.8% of job submissions, 29.1%...
Yun-Jia Zheng, Bintang Dwi Marthen, Zachary Pan et al.· 0 citations
Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on $\tau^3$-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption.
Ming-Hao Li, Bang-Yan Li, Zi-Fan Wang et al.· 0 citations
This paper discusses how CloudWeaver scopes the context of individual agent sessions with local views of cloud resources and coordinates concurrent management operations on shared cloud resources, and offers strong safety guarantees and attributable feedback in the presence of conflicting intents, while preserving conc...
Minghao Li, Ziqian Liu, Ziyu Mao et al.· arXiv.org· 0 citations
The model is a stateful computational middlebox inside a human-centered feedback loop, with network transport, model serving, and user playback jointly shaping how the interaction evolves, and is called for an AI-native real-time communication stack that resolve the joint control problem spanning communication, computa...
HalluProp, a Propagation-aware Hallucination inference framework that estimates individual agent failures and emergent system-level hallucination risks before inter-agent interaction, and effectively complements post-hoc methods, highlighting the potential of pre-hoc risk inference for building more reliable multi-agen...
Shi Lin, Chenpei Wang, Peng Qian et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.