Preprint
Aug 2026
Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms
This work compares three fusion paradigms by the artefacts they reuse and suggests that Merge when experts already exist and cheap fusion is paramount; Mix RL when training a unified model without experts, with domain proportions adjusted for cross-domain transfer; and MOPD when preserving domain-specific gains matters more than surpassing teachers or minimizing end-to-end cost.
Sicheng Wu, Kai Yang, Yuchen Cai et al.
· 0 citations