How should an embodied agent respond when a person's correction may be wrong? We formulate grounded correction arbitration as a choice among accepting, rejecting, inspecting the world, and asking the speaker. GAVA implements this interface with observation-bounded evidence, legal probes, and a one-step expected-loss ru...
Ye-Zhou Cheng, Run-Jia Du, Ze-Ming Liu et al.· 0 citations
Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty. We characterize that requirement on a fixed grid of $M$ tasks with $L$ binary paths per task under the hard budget $(M+t)K$, where each path costs at most $K$ responses or episodes. For fixed $...
Ye-Zhou Cheng, Run-Jia Du, Ze-Ming Liu et al.· 0 citations
LWCal is proposed, a CPU-only post-hoc calibrator that down-weights calibration examples whose noisy labels are contradicted by the base model's held-out probability.
Ze-Ming Liu, Hang Lyu, Jing-Tao Zhang et al.· 0 citations
SemVerBench is introduced, the first benchmark of LLM version-constraint resolution semantics across three ecosystems (npm, PEP 440, Cargo), and six frontier models are evaluated: Opus significantly outperforms all other models, and Sonnet outperforms the OpenAI models (McNemar).
This paper is the first systematic study of whether prompt-token hidden states in contemporary LLMs exhibit Gromov Hyperbolicity (GH), a distance-based measure of tree-likeness.
Zhi-Chao Yang, Yuanze Hu, Gen Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.