Preprint
Aug 2026
StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning
This work introduces StructReward, a compute-efficient framework that provides dense reinforcement signals through structured step-level reward alignment and substantially reduces the computational overhead of multimodal reinforcement learning.
Yifan Li, Ruxi Sun, Tongzhou Zhao
· 0 citations