Preprint
Jul 2026
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
This work presents LongStraw, an objective-aware, architecture-aware system for resident-state virtualization, response replay, and distributed-gradient execution that bounds the live training graph by the response suffix while reusing the expensive prompt computation across the complete GRPO group.
Changhai Zhou, Kieran Liu, Yuhua Zhou et al.
· 2 citations