AI news for builders and product teamsUpdated Oct 10, 2026, 17:01 UTC
Research news
The latest Research stories across our sources, prepared from the publishers’ own reporting.
In this topic
Newest first
OpenAI published over 700 AI-generated manuscripts claiming solutions to open mathematics problems, prompting a math blog to collect more than 100 researcher responses ranging from interest to grief. Fields Medalist Hugo Duminil-Copin said he was "paralysed."

An analysis argues that AI progress over the next few years will be driven mainly by automated engineering and efficiency gains, not by models that are fundamentally different in nature, so superintelligent behavior outside math and coding remains unlikely.

OpenAI released 722 manuscripts (now 719 after three withdrawals) of mathematical results produced by an internal frontier model, including 90 of the top 500 open problems in mathematics, organized into 372 families and posted on GitHub. The model worked an average of three hours of compute per solution after being asked to attempt about 4,000 problems, mostly from a single prompt.

Periodic Labs, launched in September 2024 by Liam Fedus and Ekin Dogus Cubuk, is building AI systems that learn from physical experiments rather than published results. The company aims to compress decades of scientific trial-and-error into months using autonomous labs.

OpenAI solved 90 of the top 500 open math problems using an average of three hours of Pro-level compute per question, which the author calls the biggest day so far in mathematics. Anthropic also released Claude Haiku 5.5 at $0.10/$0.50 per million tokens.

MIT associate professor Christina Delimitrou applies machine learning to make large-scale data centers more efficient, secure, and reliable by redesigning cloud systems, managing shared hardware resources, and streamlining server architectures. Her research aims to reduce data center power consumption by improving utilization of existing hardware.

Apple researchers introduce Normalizing Trajectory Models (NTM), which model each reverse diffusion step as a conditional normalizing flow with exact likelihood training. On text-to-image benchmarks, NTM matches or outperforms strong baselines in just four sampling steps.

A three-month randomized trial with 133 patent attorneys found AI drafting tools improved work quality for all lawyers, but only senior lawyers (7+ years) showed gains in unassisted judgment; juniors saw no average skill improvement and their scores split into more strong and more low results.

NVIDIA's Seattle Robotics Lab and Isaac team built robots to assemble GB300 tester trays, reaching over 95% success on busbar assembly with a 160-second cycle time, though short of the 124-second target.

Gary Marcus criticized OpenAI's report on a new math result as too vague to assess, noting unknown procedures and failure rates, and said the system's generalizability beyond math is unclear. The essay promises a second section with Terence Tao's take.

IEEE Spectrum and Wiley have published a sponsored white paper describing HiPHI, a 617.5-hour whole-body human motion dataset captured with optical motion capture at sub-millimeter accuracy, including 245.7 hours of human-object interaction data. It reports policies trained on the data deployed on a physical Unitree G1 humanoid robot.

OpenAI has published 372 AI-generated mathematical results on GitHub, including Lean formalizations for machine verification, with each result averaging about three hours of ChatGPT Pro compute. Twenty-five Fields Medal winners warned in an open letter that mass-producing mathematical truths could destroy fertile ground rather than bring new ideas to life.

Jake Boggan commented on Hacker News that OpenAI's math system has supposedly proven Barnette's Conjecture, a problem he worked on intermittently for 24 years.

OpenAI has released 722 manuscripts containing solutions to mathematics problems produced by an unreleased frontier model, covering 372 result families. An independent advisory group, AGMAI, says the batch includes solutions to hundreds of open questions.

Zvi Mowshowitz's newsletter covers grade inflation at Harvard, the effects of holistic admissions, standardized testing, and disability accommodations. It cites a paper by Raj Chetty, David Deming and John Friedman on Ivy-Plus admissions and a Harvard report finding grading practices failing key functions.

Since 2018 the Lincoln Laboratory Supercomputing Center has run the Lincoln AI Computing Survey, which compares commercial AI accelerators on peak performance and peak power. Its latest paper examined more than 120 accelerators, up from 57 in the first paper.

Simon Willison commented on EmbeddingGemma 2, praising its Apache 2.0 license and arguing that closed, hosted-only embedding models create risk because stored embedding vectors must be recalculated if a vendor discontinues a model. He said he prefers paying a provider to host it while retaining the option to run the open weights himself.

In a Microsoft Research Podcast episode, partner research manager Jennifer Neville discussed with applied scientist Chad Atalla how AI evaluation pushes model performance, the surprising failures found when models are tested beyond traditional benchmarks, and practical guidance for working with current AI systems.

Google Research presents five partner-led case studies using its Earth AI Population Dynamics Foundation Model (PDFM) as a geospatial foundation model for public health. PDFM's location embeddings matched or improved conventional inputs across vaccination, cardiovascular disease, dengue, postpartum depression, and cholera tasks.

OpenAI published new results on open problems in mathematics from an internal frontier model, and shared Lean proof formalizations and research details on GitHub.