Skip to content

Author

Chenxu Zhao

We have 9 of 11 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

AR-WAM: A Visual-Conditioned Agent-Ready World Action Model for Robotic Manipulation

AR-WAM is presented, a visual-conditioned, agent-ready world action model that replaces language with two complementary conditions: a visual grounding prompt and a learnable operation token dictating the atomic skill to execute, predicting scene evolution within compact latent states while decoding actions.

Yi-Cheng Jiang, Ze-Sen Gan, Xiao-Bo Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

UniOPSD: Unifying Outcome and Hindsight Feedback for Agentic Reinforcement Learning

Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy's...

Zeng-Huang Fu, Zhao-Yang Li, Qiu-Yuan Ai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SIPO: Selective-Inference Policy Optimization for Tree-Structured Agentic RL

Tree-structured reinforcement learning trains search agents by comparing alternative continuations and propagating terminal rewards to intermediate decisions. Adaptive expansion, however, creates a statistical asymmetry: an incumbent is selected using its own generation statistic, whereas fresh siblings are sampled aft...

Zeng-Huang Fu, Ning Chen, Ming-Da Jia et al. · 0 citations
Jul 2026

Self-Supervised Skill Optimization

Self-Supervised Skill Optimization (SSO) is introduced, a comparative framework that learns a reusable skill from unlabeled task instances alone and outperforms existing GT-free prompt optimizers on both closed-ended and open-ended tasks.

Siran Peng, Cui-Yu Yang, Tianyu Fu et al. · 0 citations
Jul 2026

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

WebRetriever is introduced, a large-scale benchmark encompassing 800 websites and 1,550 tasks across diverse domains, including consumer, professional, and enterprise sectors, with comprehensive coverage of user intent patterns, and NavEval (Navigation Evaluation), a novel LLM-as-Judge framework that leverages rich int...

Wei Dong, Tianyu Fu, Zhe Yu et al. · 0 citations
Jul 2026

Dive Into the Implicit Biases of Low-rank Vision-language Alignment

This work systematically characterize the implicit biases introduced by low-rank adaptation during alignment and establishes two theorems showing that low-rank alignment induces preferences for parameter subspaces with flat gradients and feature subspaces robust to perturbations, providing a principled explanation for...

Mingjia Shi, Shuo Wang, Xiaobo Wang et al. · 0 citations
Preprint Aug 2026

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

LiLa-WAM is proposed, a lightweight world-action model that reasons about the future in a compact latent space and can be trained end-to-end on a single 24GB GPU and the Visual Transition Token (VTT), a language-free task representation that encodes each task as a direction in visual feature space.

Fan Yang, Yu-Ting Su, Xiaobo Wang et al. · 9 citations
#reinforcement learning Preprint Aug 2026

CoEvoKG: Co-Evolving Knowledge Graphs with Self-Evolving Search Agents

CoEvoKG is introduced, a framework that turns a knowledge graph into both a source of verifiable training tasks and a persistent evidence memory for agent evolution, closing the loop between model self evolution and knowledge accumulation.

Zhaoyang Li, Zenghuang Fu, Qiuyuan Ai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.