Skip to content

Author

Jun-Hua Liu

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Bait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box Editing Jailbreaks

Open-weight large language models face a low-cost white-box threat from representation engineering attacks. Attackers can estimate refusal directions and search for projection-matrix edits that suppress safety alignment while preserving general capabilities, within minutes on a single GPU and without gradient-based tra...

Tian Gao, Zhi-Hui Xie, Yu-Hao Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges

C-SafeQA, a policy-grounded benchmark for response-level Chinese safety evaluation, is introduced, with substantial trade-offs between unsafe-response recall and risk-query-conditioned safe-response false positive rate.

Rui-Wei-Chun-Hui-Li-Yuan Yang, Shuang Huang, Jun-Hua Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.