Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, such evaluation of end-state performance lacks transparency into how and why agents fail in various tasks, obfuscating cr...
Zhi-Lin Wang, Shao-Kun Zhang, Yi-Fan Zhang et al.· 0 citations
Zone of Proximal Policy Optimization (ZPPO), inspired by Vygotsky's zone of proximal development, is introduced, which outperforms off/on-policy distillation and GRPO, with the largest gains at the smallest scale.
RADIO1D is introduced, which compresses images into a compact, variable-length 1D token sequence using multi-teacher knowledge distillation and an autoencoder design, delivering competitive performance on diverse multimodal benchmarks with lower computational overhead and better accuracy.
Greg Heinrich, Michael Ranzinger, Collin McCarthy et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.