Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
This work introduces Co-RL, a framework in which multiple decoupled models, sharing no parameters, are simultaneously optimized through RL using rewards derived from their peers, and shows that unsupervised reasoning can emerge through cooperative multi-agent training.
Yunhao Yang, Yuexin Bian, Yunjie Tian et al.
· 0 citations