BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning
BASIS, a critic-free post-training algorithm designed to address the tradeoff between computational efficiency and sample efficiency in value estimation and policy learning, achieves performance close to multi-rollout GRPO-type baselines and often outperforms single-rollout REINFORCE-type baselines.