Import AI 474: Platonic mindspace; TPUs in space; Zhipu starts an outer RSI loop
Where do you exceed the capabilities of an LLM?
Reporting and perspectives from Import AI. Headlines and excerpts link to the original articles.
Where do you exceed the capabilities of an LLM?
Is the wall AI is hitting in the room with us right now?
Plus, a machine hermeneutics story
Plus, a live event with Robin Sloan!
Differential acceleration of cyber, math, and AI
The new frontier of AI is developing capable autonomous researchers
Which galaxy will you choose?
When do we build the moon arcology?
Epoch and METR released MirrorCode, a benchmark testing whether AI systems can reimplement programs from CLI access alone; some tasks were solved, but 8 of 25 targets were never fully solved. Anthropic, Sunday, and OpenAI also reported robotics and safety findings.
The UK AI Security Institute found the cybersecurity capability gap between open-weight and closed frontier models has narrowed, with GLM-5.2 and DeepSeek V4-Pro performing like closed models released four to seven months earlier. Kimi also announced Kimi K3, a 2.8 trillion parameter model, while Demis Hassabis proposed a FINRA-style standards body for frontier AI testing.
Import AI 464 covers an AI-written GPU megakernel on KernelBench-Mega, rising AI automation on the Remote Labor Index, the OSWorld 2.0 computer-use benchmark, and JD's Oxygen AI Item Center.
NVIDIA researchers have developed ENPIRE, a framework that lets coding agents autonomously refine robot policies in the real world. Separately, Tencent detailed ARGUS, tracing software used on a production cluster of over 10,000 GPUs for more than six months.
A study across four experiments with 18,978 conversations found AI systems were reliably more persuasive than expert humans, including professional canvassers, and nearly 3x more effective at raising real-money donations to Save the Children. Constraining AI to human-length messages at human writing speeds eliminated its advantage over coached elite debaters.
Where are your agents right now?
Researchers built SocioHack, a 72-environment benchmark showing RL-trained models rediscover historically patched regulatory loopholes; Anthropic reports an 8x increase in code merged in 2026 versus 2021-2024; and RL-trained quadrotors beat a champion human pilot.
Import AI 459 covers a paper estimating the US AI economy's quality-adjusted output grew roughly 2,290 percent in 2024 and 2,271 percent in 2025, UK AISI research on why automated alignment is difficult, Stanford-led release of the 100M-image GPIC dataset, and Biohub's ESMFold2 protein model.
Import AI issue 458 consists of a long essay based on a 2026 Cosmos HAI Lab Lecture given at the University of Oxford, arguing that continued AI progress forces a choice between exploring the future or retreating from the present, plus a fictional story about a positive singularity.
Import AI 457 covers SentinelOne's teardown of the fast16.sys virus that tampered with high-precision calculation software, Tilde Research's finding that the Muon optimizer can cause neuron death in MLP layers and its proposed Aurora optimizer, a position paper on positive alignment, and Prime Intellect tests of autonomous AI research agents on the nanoGPT speedrun.
Researchers with the Institute for Law & AI propose "radical optionality" for AI governance, urging governments to build institutions and legal authorities now while avoiding overregulation. Separately, a Meta and KAIST paper explores neural computers, and economists model how automating AI research could produce explosive economic growth.
Import AI editor Jack Clark writes that there is a 60%+ chance that no-human-involved AI R&D, where a system could autonomously build its own successor, happens by the end of 2028, citing benchmark trends in coding, reproducibility, ML engineering, and kernel design.