Skip to content

5 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Sep 2026

Dense Is Not Enough: Hierarchical Supervision Allocation for Long-Horizon On-Policy Distillation

On-policy distillation (OPD) transfers the capabilities of a large language model to a smaller student by providing teacher supervision on the student's own rollouts. In long-horizon agentic tasks, however, uniform token-level matching can allocate supervision poorly: a large local discrepancy need not improve future b...

Yu-Hao Sun, Bin-Rui Wu, Zhuo-Er Xu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SkillLift: Learning Dense Rubrics from Sparse Oracles for Efficient Skill Evolution

LLM-based agents increasingly rely on persistent skills, i.e., reusable procedural prompts, to adapt without weight updates. Existing skill self-evolution methods directly revise skill text based on execution feedback, but each oracle evaluation requires a full agent rollout, creating a supervision bottleneck that conf...

Hao Kang, Ming Wen · 0 citations
#artificial intelligence Preprint Sep 2026

HazardAuditor: From Executable Threats to Safer Computer-Use Agents

Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing safety risks that emerge through runtime behavior rather than generated content alone. Existing guard models target static prompts and responses and are poorly suited to agent execution; existing executa...

Yun-Hao Feng, Rui-Xiao Lin, Ming Wen et al. · 0 citations
Preprint May 2026

Rewriting the Response Path: Silent Tampering and Provider-Signed Defense in BYOK LLM Agents

LLM agents convert model outputs into consequential actions, including communications, code changes, and financial transactions. Developers often trust evidence such as test results and execution logs. We identify a response path integrity gap in Bring Your Own Key configurations used by roughly 88 percent of mainstrea...

Mingyu Luo, Zihan Zhang, Zesen Liu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

UI-Venus-2 Technical Report

UI-Venus-2 is presented, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework that integrates safety-aware mechanisms to ensure controlled execution of consequential actions.

Venus Team, Zhuo-Hang Cai, Hao-Xin Chen et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.