Skip to content

Author

Shikun Zhang

We have 6 of 185 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation

LLM judges increasingly evaluate responses against fine-grained rubric checklists. When a sample requires multiple rubrics, current methods typically assess each in a separate inference call. Evaluating all rubrics in a single pass is a natural alternative with greater efficiency, but we find that it introduces rubric...

Dingyao Yu, Tong Zhang, Yutao Mou et al. · 0 citations
Preprint Aug 2026

TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models

Reward models are a bottleneck for reinforcement learning in embodied AI. Long-horizon robotic manipulation requires scalable vision feedback beyond handcrafted rewards or task-specific annotations. Existing open-source VLM reward judges like RoboReward adopt simple 1--5 trajectory progress scoring, lacking pairwise pr...

Yi-Dong Wang, Yan Zhan, Ziteng Feng et al. · 0 citations
Jul 2026

SHIFT: Self-reconstruction Harnesses Implicit Fine-grained Thinking for Retrieval

This work proposes SHIFT, a retrieval training framework based on LLMs that transfers LLMs into reasoning-efficient retrievers with residual projection and task-oriented bidirectional attention aggregation in the latent space, and alleviates the mismatch between contrastive learning and implicit reasoning using fine-gr...

Yuxiao Luo, Da Li, Mingjie Zhang et al. · 0 citations
Preprint Aug 2026

ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

Experiments reveal substantial agent vulnerabilities and show that injection timing and placement affect attack effectiveness, and ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility.

Yutao Mou, Pengfei Yang, Zhenfei Yin et al. · 0 citations
Conference Open access 2026

LeLoRA: Learnable Low-Rank Adaptation of Large Language Models

A novel Le arnable Lo w-R ank A daptation (LeLoRA) framework that utilizes dynamically learned fine-tuning strategies to facilitate the effective adaptation of LLMs and provides compelling evidence that LeLoRA consistently outperforms existing baselines in adapting LLMs.

Xiaoling Zhou, Mingjie Zhang, Zhemg Lee et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.