Preprint
Aug 2026
MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents
MemOPD records the inputs and sampled outputs of each model invocation, restores its original token positions and causal visibility, and packs the reconstructed invocations for efficient teacher scoring, which improves F1 over PPO by up to 4162%, while packing yields up to a 1.63x speedup in actor computation during training.
Zhiyuan Liu, Tinghong Ye, Chenghao Liu et al.
· 0 citations