AI #189: New Math

OpenAI released solutions to 90 of the top 500 open mathematics problems, along with many others, using an average budget of roughly three hours of Pro-level compute per question. Zvi Mowshowitz, writing in AI #189, described this as the biggest day so far in the history of mathematics and said the drop was more significant than any new AI model released that week, with further coverage planned for the next day.
Anthropic released Claude Haiku 5.5, priced at a tenth of Haiku 4.5 at $0.10/$0.50 for prompts up to 100,000 tokens, with cache writes and reads at $0.125/$0.01. Prices rise fivefold for longer prompts, so API users are advised to set automatic compaction at 100k. Anthropic recommends the model for high-volume, cost-sensitive tasks but says it underperforms Sonnet or Opus on a cost-performance basis once tasks get complex. Anthropic also cut the cost of Sonnet 5.5 cache reads by 50%.
Other items in the roundup: Jay Clayton is the new AI Czar heading a new taskforce; David Robinson resigned amid the ongoing "preference cascade," and three OpenAI safety employees were fired; Anthropic expanded its cyber verification program; a16z published its seventh Top 100 Consumer AI Apps index, which reports that only 25% of U.S. consumers use AI daily; and OpenAI rolled out watermarking within the EU only.
The roundup also cites a Taste evaluation in which Opus 5.5 reportedly has 2.3 times the experimental research taste of the best human experts, matching their score with about 17 GPU hours of experiments instead of 40. Other mentioned releases include Nano Banana 2.1 and Google DeepMind's EmbeddingGemma 2, a 740M-parameter natively multimodal open model for on-device embeddings. Zvi notes the continued wait for Gemini 4 Argon.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
The big drop of this week was not a new AI model.