Skip to content

RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling

Sep 2026 · 1 citation · 31 references
Computer Science

TL;DR

RewardVerse, a rubric-based video reward framework that introduces a dynamic rubric as an intermediate representation between the evaluation query and the scorer, mitigates scalar drift, and provides a robust and interpretable reward signal for RL in video generation.

Abstract

Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, existing video reward models often produce unstable scalar scores because they directly map complex, subjective video quality into a single score without explicit evaluation criteria. This leads to scalar drift, where the scoring scale collapses or shifts across different prompts, making the reward unreliable for RL. Drawing inspiration from professional human annotation engineering, we address this problem with RewardVerse, a rubric-based video reward framework that introduces a dynamic rubric as an intermediate representation between the evaluation query and the scorer. Instead of unconstrained direct scoring, RewardVerse first generates explicit evaluation criteria and then performs rubric-guided scoring, providing a stable semantic anchor that mitigates scalar drift. To efficiently optimize this collaborative pipeline, we propose Rubric-Guided Policy Optimization (RGPO), a two-stage training algorithm. RGPO first warms up the scorer using self-evolving seed rubrics and then jointly optimizes the rubric generator to produce query-adaptive evaluation criteria while continuously aligning the scorer with human ratings. Extensive experiments on the 16-dimensional EvalVerse benchmark and external datasets demonstrate that RewardVerse mitigates scalar drift, achieves state-of-the-art performance on both pointwise and pairwise evaluation, and provides a robust and interpretable reward signal for RL in video generation.

View source

Similar papers

Preprint Aug 2026

RubricRM: Generative Reward Modeling via Dynamic Rubrics for Image Generation and Editing

Reward models play an essential role in aligning visual generative models, yet most existing visual reward models use a single scalar score or rely on fixed criteria that cannot adapt to different instructions. This limits both interpretability and task sensitivity, especially for text-to-image generation and instructi...

Zi-Jian Kan, Wei Wang, Long Luo et al. · 0 citations
Preprint Sep 2026

Think Before You Score: Thinking Reward Model for Visual Generation

Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case. We introduce Think Before You Score, a paradigm that expli...

Xue-Yuan Bai, Zhen-Chen Tang, Yang Shi et al. · 0 citations
Preprint Sep 2026

HiRE: Hindsight Reward Editing for Policy Finetuning

Pre-trained robot policies always require finetuning to adapt to specific environments. Reinforcement Learning (RL) offers high performance potential because it improves action optimality rather than simply mimicking data. However, such potential depends heavily on reward quality. Sparse rewards lack process feedback,...

Haoyi Niu, Zhen Han, Yu-Feng Ji et al. · 0 citations

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards

This work proposes Step-Aware Annealing (SAA), a plug-and-play reward sharpening mechanism that progressively increases reward curvature during training, amplifying subtle quality differences among high-scoring samples while preserving stability in early learning.

Unknown authors · 0 citations
2025

LaRes: Evolutionary Reinforcement Learning with LLM-based Adaptive Reward Search

This work proposes LaRes, a novel hybrid framework that achieves efficient policy learning through reward function search by leveraging large language models to generate the reward function population, guiding RL in policy learning.

Pengyi Li, Hongyao Tang, Jinbin Qiao et al. · 6 citations
Preprint Aug 2026

TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models

Reward models are a bottleneck for reinforcement learning in embodied AI. Long-horizon robotic manipulation requires scalable vision feedback beyond handcrafted rewards or task-specific annotations. Existing open-source VLM reward judges like RoboReward adopt simple 1--5 trajectory progress scoring, lacking pairwise pr...

Yi-Dong Wang, Yan Zhan, Ziteng Feng et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.