AivexaNewsSearch
AI news for builders and product teamsChecked every hour
Import AIResearch

Import AI 455: AI systems are about to start building themselves.

Collected Sep 30, 2026

In an essay for the Import AI newsletter, Jack Clark writes that he has "reluctantly" concluded there is a likely chance (60%+) that no-human-involved AI R&D — defined as an AI system powerful enough to plausibly autonomously build its own successor — happens by the end of 2028. He states he does not expect this in 2026, but thinks an example of a model end-to-end training its successor could appear within a year or two, certainly as a proof-of-concept at the non-frontier model stage, while frontier models may be harder because they are more expensive and involve many humans.

He cites public information including papers on arXiv, bioRxiv and NBER and products deployed by frontier companies, and warns that benchmarks have idiosyncratic flaws while arguing the aggregate trend matters.

On coding, he notes SWE-Bench's best score went from about 2% at its late-2023 launch with Claude 2 to 93.9% with Claude Mythos Preview, and cites METR's time-horizon measure: about 30 seconds in 2022 with GPT-3.5, 4 minutes in 2023 with GPT-4, 40 minutes in 2024 with o1, roughly 6 hours in 2025 with GPT 5.2 (High), and about 12 hours in 2026 with Opus 4.6. He says METR's Ajeya Cotra considers it not unreasonable to expect about 100-hour tasks by the end of 2026.

On scientific skills, he reports CORE-Bench's hardest tasks rose from about 21.5% by a GPT-4o CORE-Agent to 95.5% with Opus 4.5, which a benchmark author declared "solved" in December 2025; MLE-Bench's top system rose from 16.9% in October 2024 to 64.4% in February 2026. On PostTrainBench, he says top systems score 25%-28% as of April versus a human score of 51%. On Anthropic's CPU-only small language model training task, mean speedups rose from 2.9x (Opus 4, May 2025) to 52x (Claude Mythos Preview, April 2026), where a 4x speedup is expected to take a human researcher 4 to 8 hours.

He also cites an Anthropic proof-of-concept in automated alignment research, AI systems supervising sub-agents in products like Claude Code or OpenCode, and asks whether AI research is more like discovering general relativity or Lego, arguing AI cannot yet invent radical new ideas but may not need to.

Read at Import AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

The first step towards recursive self improvement