Skip to content

Author

Bing-Chen Zhao

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Review May 2026

SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents

SpecBench is introduced, a benchmark comprising 30 systems-level programming tasks ranging from short horizon tasks like building a JSON parser to ultra long horizon tasks like building an entire OS kernel from scratch, which offers a principled testbed for measuring whether coding agents build genuine working systems...

Bing-Chen Zhao, Dhruv Srikanth, Yuxiang Wu et al. · 25 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.