Skip to content

Nereus: Adaptive Parallelism for LLM Post-Training

Sep 2026 · 1 citation · 58 references
Computer Science

TL;DR

Nereus is a cost-aware runtime that adapts RL post-training jobs into efficient execution plans and executes a transition using a memory-feasible global plan and admits the transition using a cost model calibrated against the running job.

Abstract

Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a consequence, an execution plan that was initially suitable can then become slow or even infeasible over time. However, adapting a job whose models share GPUs entails significant challenges: deciding whether a new plan is worth the transition cost, reusing the job's distributed state, and coordinating GPU transfers across models and stages. Nereus targets these challenges as a cost-aware runtime that adapts RL post-training jobs into efficient execution plans. Its low-overhead controller selects a memory-feasible global plan and admits the transition using a cost model calibrated against the running job. To estimate and execute a transition, Nereus represents the distributed state of each replica of a model-stage (one model in one stage) as an Elastic Model Unit. It then employs a global transition graph to order the transformations and GPU transfers of these units. In a trace built from real data, online TP/PP adaptation reduces average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling. In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time. Nereus improves end-to-end 8B PPO throughput by 2.14--7.27$\times$ over OpenRLHF and by 1.10--1.47$\times$ over Verl across diverse clusters.

View source

Similar papers

Preprint Aug 2026

Performance Foundations of Parallel&Distributed Reasoning Language Models

This work systematize the RL-for-LLM paradigm and provides a compute-centric analysis of prominent post-training algorithmic frameworks: Proximal Policy Optimization (PPO), Group Relative Policy Optimization (GRPO), as well as their variants, and develops a taxonomy of intra- and inter-model parallelism strategies for...

Maciej Besta, L. Schmidt, Lara Nonino et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Frontier Learning: Training LLM Reasoners at the Edge of Capability

Across several reasoning tasks and model families, the proposed frontier learning approach consistently achieves higher relative gains over fixed-pool baselines, demonstrating that effective post-training requires not only selecting useful problems, but continually generating them at the edge of capability.

Robin Faro, S. Ramesh, Ilija Bogunovic et al. · 0 citations
#machine learning Preprint Sep 2026

QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents

QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs.

Wei-Qi Wang, Yu-Xin Zhou, Mou-Xiang Chen et al. · 0 citations
#natural language process... Preprint Oct 2026

LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches

Reinforcement learning has greatly advanced the capabilities of large language models, but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank gradient sketches. These compact representations su...

Shao-Kun Zhang, Yi-Fan Zhang, Jian Hu et al. · 0 citations
#artificial intelligence Preprint Oct 2026

ThunderSyncRL: Lossless Acceleration of Agentic Reinforcement Learning

Language models are moving beyond generating answers to pursuing long-horizon goals in interactive environments. Post-training these agents requires long, heterogeneous trajectories, and synchronous systems leave learner engines idle until rollout and verification finish. To squeeze out these pipeline bubbles, asynchro...

Seil Kang, Hangoo Kang, Tarun Suresh et al. · 0 citations
#machine learning Preprint Sep 2026

EBRL: Asynchronous Embodied RL by Multi-Grained Resource Management

Embodied reinforcement learning (RL) improves model capabilities with a pipeline of environment simulation, action generation, and model updates. These stages show heterogeneous CPU and GPU demands, making efficient resource utilization difficult. Recent systems overlap rollout (simulation and generation) with training...

Liang Mi, Wei-Jun Wang, Bo-Wen Gao et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.