Skip to content

Author

Semih Yavuz

We have 2 of 72 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Learning Generalizable Behaviors for Terminal Agents

River, a simple training recipe that improves reward quality by filtering low-quality environments and augmenting outcome rewards with process-level behavior regularization is proposed, which achieves the best performance among evaluated open-source RL-trained 8B models across four terminal-agent benchmarks.

Yi-Fan Yao, Bo Pang, Xuan-Phi Nguyen et al. · 0 citations
Preprint Jul 2026

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models

Procedural Memory Distillation is proposed, which converts crossepisode signals into reusable procedural memory and distills it into the policy's weights during training, yielding a memory-free model at inference.

Ye Liu, Srijan Bansal, Bo Pang et al. · 2 citations