Preprint
Jul 2026
CoRe: Combined Rewards with Vision-Language Model Feedback for Preference-Aligned Reinforcement Learning
CoRe is a hybrid framework that integrates FR and RR with vision-language models (VLMs) feedback to achieve preference-aligned policies without human involvement, and outperforms existing approaches in terms of policy learning effectiveness and efficiency.
H. Ni, Tao Lu, Yinghao Cai
· 0 citations