Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing
Import AI Jack Clark
Researchers at King's College London, Fudan University, and The Alan Turing Institute created SocioHack, a benchmark with 72 simulated environments to test whether AI systems can discover loopholes in institutional rules while remaining technically compliant. In historical environments reconstructed from real regulations, reinforcement learning-enabled language models rediscovered previously patched exploits with 61.25% recall and 90.85% precision, including strategies for ocean mining rights and credit card rewards. As AI systems become more capable at both quantitative and qualitative reasoning, automated exploitation of bureaucratic gaps could create widespread institutional vulnerabilities.
Why it matters
When will markets price the singularity?
Related stories
Import AI 469: Science AI; RSI simulator; and Zuck's technological pessimism
Import AI · 1 month ago ·
18
Import AI 468: 23 RSI ideas; PostTrainBench+; and how trust and transparency interplay with AI racing
Import AI · 1 month ago ·
32
Reward Hacking in Reinforcement Learning
Lil'Log · 1 year ago ·
6