Skip to content

Author

Jian-She Li

We have 6 of 13 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Review Sep 2026

BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs

BENCHCOMPASS is introduced, a payment-domain benchmark whose construction pipeline builds scenario-grounded tasks from typed evidence packs, applies LLM-based quality checks, creates task-input attack variants, and reserves final item admission for domain experts.

Si-Jie Dong, Wei-Feng Ren, Xuan-Wei Hu et al. · 0 citations
#natural language process... Preprint Sep 2026

Grounded Skill Synthesis from Code at Scale for Agentic Intelligence

Reusable skills give agents transferable procedural knowledge, making scalable acquisition essential for extending agents beyond prior experience. Existing methods face two limitations: trajectory-based synthesis requires interactions with specific environments, while document-derived skills may lack executable evidenc...

Yong-Qi Tong, Pan Wang, Hang Wang et al. · 0 citations
Preprint Aug 2026

ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction

The proposed ARC (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization, is proposed.

Yong-Qi Tong, Tan Li Hui Faith, C. Marcus et al. · 0 citations
Preprint Aug 2026

Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning

ACA-RL supports a new mission for NLP evaluation: measuring whether models can recognize when a task is underdetermined and handle uncertainty, not only whether they can answer fully specified questions.

Yong-Qi Tong, Zhenyu Zhang, Zimou Liu et al. · 0 citations
Preprint Aug 2026

STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment

This work proposes \methodname, a stability-guided active-set controller for controlled objective admission, a stability-guided active-set controller for controlled objective admission in reward-vector RLHF, which positions objective-entry timing as a concrete control variable in reward-vector RLHF.

Yong-Qi Tong, Z. Zhang, Ruirui Wang et al. · 1 citation
Preprint Aug 2026

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction

DARC is proposed, a diagnosis-guided recovery harness that profiles task-family failure modes, prunes mismatched interventions from a shared recovery library, and freezes a verifier-selected success-cost policy for deployment, providing a practical route toward more reliable agents in domains where compiler-like feedba...

Pan Wang, Yihao Hu, Hang Wang et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.