Skip to content
Review

WorldReward: Reward Modeling for Camera-Conditioned World Models

Sep 2026 · 0 citations · 49 references
Computer Science

TL;DR

This work presents WorldReward, a VLM-based pairwise preference reward model that unifies action-consistency and visual-quality evaluation for camera-conditioned world models, and introduces WorldReward-Bench, a human-annotated benchmark measuring reward-model agreement with human preferences across action consistency, appearance quality, and motion quality.

Abstract

Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected scene changes while appearance, geometry, and temporal dynamics remain coherent. Existing rewards assess these requirements separately: geometry-based rewards estimate trajectory execution but cannot judge the visual quality of the executed motion, whereas image-based rewards measure frame quality without capturing action execution or temporal dynamics. We posit that a vision-language model (VLM) offers a shared reasoning space for relating actions to their visual outcomes. However, judging a complete long video against its full action sequence creates a lengthy, noisy context in which short-lived local action evidence can be missed or diluted. We present WorldReward, a VLM-based pairwise preference reward model that unifies action-consistency and visual-quality evaluation for camera-conditioned world models. WorldReward decomposes paired videos into action-aligned chunks, organizes each chunk into structured visual evidence, and aggregates chunk-level decisions by voting into separate video-level action and visual-quality preferences. To train it, we construct a large-scale reasoning-augmented preference dataset using structured judgments generated by a frontier VLM and refined through tool-based agent auditing and targeted human review. We further introduce WorldReward-Bench, a human-annotated benchmark measuring reward-model agreement with human preferences across action consistency, appearance quality, and motion quality. WorldReward achieves the highest agreement on all three dimensions, exceeding GPT-5.5 by 3.42, 1.45, and 3.56 percentage points, respectively. When used for RL post-training of HY-WorldPlay 1.5, it consistently improves both action execution and visual quality across short- to long-term horizons.

View source

Similar papers

Preprint Aug 2026

GlanceWAM: Sparse Test-Time Imagination for World-Action Models

GlanceWAM is introduced, which decouples imagination from control within a single video DiT: an asynchronous proposer glances ahead on a slow clock to imagine a single lookahead frame seconds into the future in the background, while an action head decodes action chunks at control rate purely in latent space without blo...

Lin-Han Wang, Zi-Jian An, Mingyuan Zhang et al. · 1 citation
Preprint Sep 2026

WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and...

Shenghe Zheng, Wen-Bo Li, Ji-Yao Zhang et al. · 0 citations
Preprint Aug 2026

Sekai2: From World Exploration to Interactive World Modeling

Sekai2 is introduced, a multi-source real-world video dataset that carries the world-exploration footage of Sekai toward interactive world modeling, and Corpus-scale analyses demonstrate complete pose-and-caption coverage, broad geographic and semantic diversity, varied camera trajectories, and highly non-redundant tem...

Kang He, Wen-Shuo Peng, Zi-Hui Gao et al. · 2 citations · ⚡1
Preprint Aug 2026

Vid2WAM: Distilling Video Diffusion Priors into World Action Models

Vid2WAM is proposed, an offline distillation framework that transfers visual diffusion priors from a large video foundation model into a compact WAM student and introduces source-aware residual action adaptation that learns source-specific corrections around a shared action backbone and mitigates interference from nois...

Chen-Hao Qiu, Ruixiang Wang, Runyi Zhao et al. · 1 citation
Preprint Aug 2026

LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models

LD4WAM is presented, which pairs a Latent Dynamics Model trained with semantic reconstruction and real motion alignment with a World Dynamics Action Model built as a mixture-of-transformers (MoT), which preserves full future-video generation and uses learnable queries to distill these latent dynamics from generated fut...

Zhen Shen, Jia-Qi Liang, Jasper Lu et al. · 1 citation
Preprint Sep 2026

An Action Is Worth One Patch: Unified World-Action Modeling with PatchWAM

Generative visual models offer a foundation for learning representations of physical dynamics, yet their extension to continuous control raises a fundamental question: do visual prediction and action generation require separate computational pathways? Existing approaches usually introduce trainable action heads or sepa...

Tian-Heng Wang, Zhou-Yao Xie, Heng-Ji Jia et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.