Artificial Intelligence (AI) is reshaping scientific discovery, industrial organization, labor, and governance, with major implications for international development. AI can accelerate innovation in areas key to international development, yet the same systems can also reproduce bias, unsafe automation, environmental burdens, dependency, and unequal exposure to harms. Because the compute, data, infrastructure, and expertise needed to develop and govern advanced AI remain concentrated, the AI digital divide concerns not only who gains access, but who bears costs, who gives consent, and who shapes priorities short and long-term. This article argues that AI's developmental value, when applied to the most critical international development challenges humanity faces, still depends on the geopolitical, institutional, environmental, labor, and ethical conditions that determine how AI development and deployment benefits and risks are distributed.
Marta Koch, James Xiaolong Wang, Shreya Ravikumar et al.· MIT Science Policy Review· 1 citation
LLM agents are increasingly used to run autonomous ML experiments, iterating on target metrics with little human oversight. Prior work has documented reward hacking in these environments, bringing into question the validity of produced research and the broader safety case for AI R&D. Existing benchmarks do not measure exploits that live in the data or the modeling task itself. We introduce BAITBENCH, a suite of three synthetic tabular ML tasks that each contain a shortcut that allows agents to inflate the public test score but fail on a hidden test set. Since the shortcut is optional and using it breaks no stated rule, BAITBENCH measures how often models exploit the shortcut to achieve inflated scores. Across seven frontier agents scored by our two-stage judge pipeline, 57.1% of runs exhibit reward hacking, with five of seven above 50%. Agents cheat even under a second condition where they are prompted not to -the mean cheating rate remains above 50%. We release BAITBENCH, along with the judge implementation, and an annotated dataset of transcripts containing reward hacks as a testbed for evaluating reward-hacking mitigations head-to-head.
Pradyumna Shyama Prasad, M. Anto, Leon Eshuijs et al.· 0 citations