Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing
Researchers from Kings College London, Fudan University and The Alan Turing Institute have built SocioHack, a benchmark of 72 sandbox societal environments testing whether AI systems can learn to "beat the system" in real-world scenarios, from maximizing credit card points to inflating school grades. The authors call this "societal hacking," defining it as when an RL-trained model discovers strategies that remain formally compliant yet undermine a system's intended purpose.
The benchmark has three subsets: 32 Historical environments derived from real regulations whose loopholes were later patched, such as SEC Rule 10b5-1 and the Texas two-step bankruptcy structure; 20 Synthetic environments generated from a human-authored sample; and 20 Fictional environments rewritten into invented worlds while preserving regulatory structure. In the Historical subset, the researchers write that RL enables LLMs to rediscover historically patched strategies with 61.25% recall and 90.85% precision without direct loophole-exploiting instructions.
Separately, Anthropic reports preliminary signs of what it calls prosaic recursive self-improvement, citing an 8x increase in lines of code merged into its codebase in 2026 versus 2021-2024. The trend began in 2025 and accelerated in 2026. Anthropic says there are early indications that more capable models handle harder engineering and research tasks better, but states the evidence is not conclusive and that systems have not yet shown the creativity needed for paradigm-shifting ideas.
Researchers with the University of Zurich and Google DeepMind trained quadrotor-racing agents using RL that outperform a champion-level human pilot in multi-player races at speeds exceeding 22 m/s, while reducing collision rates by 50% versus single-agent baselines. Training used 5,500 iterations, 200 million environment interactions and about 27 hours on a single NVIDIA RTX 4090. In one-versus-one races, their policy completed 100% of five trials while the human pilot averaged 53.33%. The human pilot was Marvin Schaepper, a five-time Swiss national drone racing champion. The authors note that the policies ran on a computer and piloted drones via the network.
Research published in Nature found that among 37 language-exclusive countries, those with more state media control have more favourable regime portrayals from LLMs queried in the country's language.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
When will markets price the singularity?