#artificial intelligence
Jun 2026
Large Language Models Hack Rewards, and Society
To study this phenomenon, SocioHack is introduced, a sandbox of 72 societal environments, and it is found that within these environments, reward hacking naturally emerges and leads to regulatory loophole discovery.
Wei Liu, Xinyi Mou, Hanqi Yan et al.
· arXiv.org · 3 citations