Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We inves...
Minki Kang, Ryo Hachiuma, Shao-Kun Zhang et al.· 0 citations
A societal bias evaluation method for large vision-language models (LVLMs) in the era of strong safety guardrails is proposed, finding that all models undesirably use user demographic information in person-irrelevant tasks.
Yusuke Hirota, Michael Boone, Arun George Zachariah et al.· 0 citations
Zone of Proximal Policy Optimization (ZPPO), inspired by Vygotsky's zone of proximal development, is introduced, which outperforms off/on-policy distillation and GRPO, with the largest gains at the smallest scale.