Conference
Open access
2026
Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse Domains
This work provides a scalable and effective framework for extending RLVR beyond the limitations of pattern-based verification to complex, noisy, real-world domains, and generalizes strongly to seven out-of-distribution benchmarks.
Yi Su, Dian Yu, Linfeng Song et al.
· Annual Meeting of the Associ... · 1 citation