Skip to content

Author

Jin-He Bi

We have 5 of 27 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

ReflectRL is proposed, a lightweight plug-and-play framework that learns from Golden Negative Trajectories during on-policy training, and first uses these trajectories to elicit Reflective Reasoning, then applies Reflective-to-Direct Policy Transition to transfer the acquired reasoning behavior back to Direct Reasoning...

Jin-He Bi, Chennan Zhou, Zeng-Jie Jin et al. · 2 citations

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence

This survey reviews Multimodal Code Intelligence, covering systems that generate, edit, refine, or reason with code under visually grounded inputs and outputs and organizes benchmarks and methods into four domains: Graphical User Interface, Scientific Visualization, Structured Graphics, and Frontier Tasks and Framework...

Xuanle Zhao, Qiushi Sun, Jingyu Xiao et al. · 5 citations
Preprint Aug 2026

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

OPD-V is introduced, a visual OPSD paradigm that instantiates privileged information through the Positive Teacher and Negative Teacher that consistently improves reasoning performance while reducing training cost.

Aniri, Jinhe Bi, Peng Liao et al. · 2 citations

PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

Empirically, PRISM reduces the end-to-end time for data selection and model tuning to just 30% of conventional pipelines, and achieves this efficiency while simultaneously enhancing performance, surpassing models fine-tuned on the full dataset across eight multimodal and three language understanding benchmarks.

Jinhe Bi, Yifan Wang, Danqi Yan et al. · 73 citations · ⚡4
Preprint Jul 2026

MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution

MetaSkill-Evolve is introduced, a two-timescale framework that makes agentic skill improvement recursive and outperforms no-skill, static-skill, and single-level evolution baselines on three agentic benchmarks, improving held-out test accuracy over the raw backbone by +23.54, +16.09, and +1.92 points respectively.

Zefeng Wang, Minxi Yan, Jinhe Bi et al. · 4 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.