SocialRL: Refining LLMs'Social Intelligence through Multi-turn Reinforcement Learning and Reward Design
This work proposes SocialRL, a multi-turn reinforcement learning framework using PPO that propagates delayed outcome rewards back to each turn, enabling long-horizon planning and demonstrates the effectiveness of SocialRL across synthetic and real social scenes, as well as standard and challenging social scenarios.