These results establish scope matching as a complementary control for persistent agent memory: certification determines whether an edit is supported, while retrieval scope determines where that evidence authorizes its use.
Ye-Zhou Cheng, Run-Jia Du, Ze-Ming Liu et al.· 0 citations
Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty. We characterize that requirement on a fixed grid of $M$ tasks with $L$ binary paths per task under the hard budget $(M+t)K$, where each path costs at most $K$ responses or episodes. For fixed $...
Ye-Zhou Cheng, Run-Jia Du, Ze-Ming Liu et al.· 0 citations
SemVerBench is introduced, the first benchmark of LLM version-constraint resolution semantics across three ecosystems (npm, PEP 440, Cargo), and six frontier models are evaluated: Opus significantly outperforms all other models, and Sonnet outperforms the OpenAI models (McNemar).
Qi-Bai Chen, Ze-Ming Liu· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.