Skip to content

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability

Reliable refusal of harmful requests is essential to the safe deployment of language models. Because excessive eagerness to please users may undermine existing refusal capabilities, reducing sycophancy offers a potential route to stronger refusal beyond the harmful scenarios covered by safety training. We investigate t...

Xu Wang, Di-Fan Zou, Xuan-Sheng Wu · 0 citations
Preprint Aug 2026

Mismatch Matters: On-Policy Distillation Beyond Token Agreement

On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repetitive loops to achieve near-perfect token agreement with the teacher despite globally flawed responses. We therefore shift our focus from ag...

Zichao Yu, Chengzhi Yu, Shengze Xu et al. · 1 citation
#artificial intelligence Preprint Aug 2026

Enhancing SAE-based Steering via Neighbor Integrated Feature Selection

This paper proposesNeighbor Integrated Feature Selection (NIFS), a plug-and-play strategy that leverages representation similarity to improve feature selection for steering and demonstrates consistent performance gains over conventional top-$k$ selection.

Yutang Liu, Xu Wang, Di-Fan Zou · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.