Skip to content

Author

Xunliang Cai

We have 5 of 40 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Review Sep 2026

LongCat-DeepResearch Technical Report

We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports. The workflow separates global planning from detailed investigation and coordinates revision at the section level. Multiple planning agents f...

Mei Zhu, Yue-Ya Xu, Wan-Li Wu et al. · 0 citations
Jul 2026

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

SA is proposed, a Stable Advantage Fusion framework that avoids entropy collapse and consistently outperforms fixed-coefficient GRPO+OPD fusion, improving the aggregate score by 0.70% across all six model-domain settings while achieving more stable training.

Yifan Ding, Xin Wei, Yoshua Y. Li et al. · 3 citations
Preprint Aug 2026

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

A systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks is presented.

Yi-Wei Li, Wanli Yang, He-Xiang Tan et al. · 1 citation
Jul 2026

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

ClawTrack is presented, a dual-assessment benchmark that simultaneously measures what an agent achieves (Task Score) and how it achieves it (Process Score) and finds that process scores effectively attribute success and failure to specific reasoning dimensions, filtering lucky passes invisible to outcome-only evaluatio...

Xingjian Wu, Xuhan Zhu, Xing-Chen Liu et al. · 0 citations
Jul 2026

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation

Contrastive Reinforced Policy Optimization (CRPO) is introduced, which reformulates agentic OPSD from a contrastive learning perspective, and conducts group-wise contrast to preserve reliable, fine-grained optimization signals.

Xingjian Wu, Junlin Liu, Xing-Chen Liu et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.