Long-horizon large language model (LLM) agents commonly retain their complete interaction history until compaction is triggered at a predefined threshold. We study Continuous Context Management (CCM), which performs compaction at every turn to prevent interaction history from accumulating in the active prompt. At each...
William Hoy, Jing-Xuan Fan, Nurcin Celik et al.· 0 citations
Reliable Enterprise Agent Deployment (READY), a framework for qualifying AI agents for deployment on enterprise workflows, and provides a basis for comparing agent systems, setting oversight requirements, and making evidence-based deployment decisions.
Veronica Chatrath, Bryan Zhu, Jingxuan Fan et al.· 0 citations
CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework.
Veronica Chatrath, Bryan Zhu, George Pu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.