What is the best compromise in a situation where different people value different things? The most common method for answering this question is to look at all the options, add up the utility per person associated with each, and pick the option with the largest sum. This approach seems like the obvious, theory-neutral s...
Jared Moore, Ye-Jin Choi, Sydney Levine· Cognition· 0 citations
Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback...
Fang Wu, Dan-Lei Xing, Yan-Jie Huang et al.· 0 citations
Zone of Proximal Policy Optimization (ZPPO), inspired by Vygotsky's zone of proximal development, is introduced, which outperforms off/on-policy distillation and GRPO, with the largest gains at the smallest scale.