Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and s...
Chia-yuan Chang, Ren-Yuan Cheng, Rui Feng et al.· 0 citations
Search agents are usually trained under a single harness. But once an agent is deployed in a real application, its harness is frequently updated (e.g., a rewritten system prompt) to fit production needs. This exposes a fragility of post-trained agents: because a learned behavior is entangled with its training harness,...
Xin-Lu Zhang, Ying-Chun Lin, Zhi-Han Zhang et al.· 0 citations
This work demonstrates that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training and introduces Agent as Policy (AGP), which places task planning and execution under the agent's control.
Meng-Zhao Jia, Yang Lin, Xi-Xin Zhang et al.· 7 citations· ⚡1
QUBRIC, a framework that co-designs queries and rubrics can make rubric-based RL a practical complement to RLVR beyond strictly verifiable tasks, provides evidence that co-designing queries and rubrics can make rubrics a practical complement to RLVR beyond strictly verifiable tasks.