Enterprises deploying AI for supply chain decisions commonly default to the largest available language model, a procurement heuristic that neglects both empirical performance and environmental cost. We benchmark six large language models across 520 supply chain tasks, simultaneously measuring decision quality and estim...
Code review is a key quality checkpoint between AI-generated code and production and reviewer adaptation is detectable in latent distributional structure rather than in classical lexical metrics, and language shifts follow rather than precede changes in approval behavior.
Multimodal clinical decision-making requires reliable reasoning over heterogeneous evidence from electronic health records, medical images, and physiological signals. Existing models typically map these inputs directly to diagnoses without explicitly assessing evidence sufficiency, tool-use requirements, or diagnostic...
Budget-aware evaluation for Active RAG evaluation is studied by recasting active retrieval as utility estimation, where retrieval is valuable only through its marginal correctness change over a no-retrieval answer.
Pinyan Qian, Su Wang, Chong Peng et al.· arXiv.org· 3 citations· ⚡1
Multi-agent applications delegate work across independently operated deployers. After an incident, a verifier must answer two questions: which deployer released the reported bytes, and whether each cross-deployer edge was authorized. Credentials establish who may act, but need not bind them to later output bytes or pro...
VERA (Verifiable Edge Revocation for Agents), a verifier-checkable revocation contract and API emitted by agent-runtime adapters as signed evidence, is introduced and schema portability on A2A, AutoGen, and CrewAI artifacts is validated.
This work introduces the Counterfactual Fabrication Lab, a deterministic micro-lab where the correct action is known: do nothing, and presents the Counterfactual Fabrication Lab for measuring fabricated failures in self-improving agent harnesses.
Su Wang, Pinyan Qian, Yifan Lin et al.· arXiv.org· 6 citations· ⚡2
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.