Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model backbone with a harness for planning, execution, memory, and verification, but this architecture still leaves domain-specific know-how outside the agent. We call this missing layer operational knowledge,...
Jianlyu Chen, Yuyang Hu, Hong-Jin Qian et al.· 1 citation
LOFA combines reinforcement learning over verifiable purchase outcomes with feedback-aware on-policy distillation, which identifies users' in-dialogue directives and converts them into dense token-level supervision, which captures both collaborative behavioral patterns and user-specific preferences.
Haobo Zhang, Kelong Mao, Sulong Xu et al.· 0 citations
This work introduces AREX, a family of Recursively Self-Improving (RSI) deep research agents that substantially outperforms comparable-scale baselines and remains competitive with models using substantially more activated parameters.
Shuqi Lu, Chaofan Li, Kun Luo et al.· arXiv.org· 2 citations· ⚡1
On-Policy Omni Distillation (OPOD), which consolidates text, image, and audio teachers into one omni model, and surpasses the base model and pooled RL training on all twelve benchmarks, and ranks first or second on eleven even when the teachers are included.
Tong Zhao, Yuyang Hu, Reed Li et al.· arXiv.org· 0 citations
This work proposes WebSwarm, a progressive recursive delegation framework that jointly constructs task decomposition, recursive expansion, and agent collaboration during inference and consistently outperforms single-agent and multi-agent baselines on deep, wide, and interleaved deep-and-wide tasks.
The Douyin Multimodal Embedding (DME), a model trained in two stages to combine both strengths, is presented, a model trained in two stages to combine both strengths.
Hao-Nan Chen, Chu Li, Zhi-Cheng Wang et al.· 1 citation
Experimental results demonstrate that CLAIM can learn stable and generalizable clarification strategies without relying on manually labeled data, offering a low-cost and robust solution for proactive understanding in real-world open-domain interactions with LLMs.
Experiments show that HiRA significantly outperforms state-of-the-art RAG and agent-based systems, highlighting the effectiveness of decoupled planning and execution for multi-step information seeking tasks.
Jiajie Jin, Xiaoxi Li, Yuyao Zhang et al.· Annual International ACM SIG...· 0 citations
This work introduces SearchOS, a system-level multi-agent framework that turns fragile, implicit search progress into explicit, persistent, and shared state, and introduces a Search Tool Middleware Harness that intercepts model and tool interactions to record grounded evidence and react to stalls or budget exhaustion.