Skip to content

Author

Congjing Ran

We have 2 of 29 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Harness-agnostic detection and immunization of reward hacking in self-evolving language models

This work introduces HackProbe, a monitor that attaches to an arbitrary self-evolving loop through two black-box hooks, with no access to weights or activations, and proves a detectability bound that converts a target error rate into an explicit probe-size budget, and delimit what probe rotation does and does not buy.

Rong-Xin Yang, Yang Liu, Shang Luo et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.