Preprint
Aug 2026
Performance Foundations of Parallel&Distributed Reasoning Language Models
This work systematize the RL-for-LLM paradigm and provides a compute-centric analysis of prominent post-training algorithmic frameworks: Proximal Policy Optimization (PPO), Group Relative Policy Optimization (GRPO), as well as their variants, and develops a taxonomy of intra- and inter-model parallelism strategies for RL-for-LLMs.
Maciej Besta, Leonard Schmidt, Lara Nonino et al.
· 0 citations