AI news for builders and product teamsUpdated Oct 10, 2026, 20:01 UTC
Research news
The latest Research stories across our sources, prepared from the publishers’ own reporting.
In this topic
Newest first
OpenAI said on September 8 that an unreleased internal model resolved the Navier-Stokes existence and smoothness Millennium Prize problem, a claim the Clay Mathematics Institute still lists as unsolved. The announcement drew a credit dispute with NYU mathematician Tristan Buckmaster, which OpenAI denied.

NVIDIA said TensorRT Edge-LLM ran Qwen3.6-27B on a single Jetson AGX Thor Developer Kit in the MLPerf Inference v6.1 Edge Agentic benchmark at 52.33 tokens per second, completing 1,007 turns in 24 minutes 36 seconds. NVIDIA reported this as 6.4x faster than the llama.cpp reference run of 2 hours 37 minutes.

PyTorch published a technical report on extending FlashAttention-4 with end-to-end MXFP8 block-scaled attention for Blackwell GPUs, reporting up to 1.6x forward and 1.52x backward gains over BF16. The code is open sourced in Meta's ads_model_kernel_library repository.

MIT and collaborating researchers developed xvr, an AI technique that matches intraoperative 2D X-rays with a patient's preoperative 3D CT or MRI scan in seconds with sub-millimeter precision. A paper on the method appears in Nature.

MIT political scientist Naoki Egami, who joined the Department of Political Science as an associate professor with tenure in 2025, studies research methodology, including external validity and errors introduced by AI tools in social science data. He has received the Society for Political Methodology's Emerging Scholar Award and several American Political Science Association best paper awards.

Good Start Labs, spun out of Every with $3.6 million in funding, trains AI models on games as verifiable training environments. Alex Duffy said a 1830 railroad game experiment found only a multi-turn terminal-agent design improved performance on the Finance-Agent benchmark, while Diplomacy fine-tuning improved customer support and industrial operations benchmarks.

Google says its technologies now support more than 300 languages spoken by 7 billion people, 86% of the global population, and released AI & Economy ATLAS insights on how people use AI worldwide. It also detailed recent science work including AlphaGenome Atlas, WeatherNext 3, and a Planetary Prediction Engine.

Google has launched an open-access interactive experience for its AI & Economy ATLAS data, alongside new research with Google DeepMind and MIT FutureTech showing nearly half of surveyed scientists use AI daily and save almost seven hours a week.

Andon Labs, a San Francisco AI safety company, runs real businesses with AI agents as managers to test how much real-world responsibility agents can handle. Its experiments, including Andon Market and Andon Café, serve as testbeds for evaluations developed with frontier AI labs.

MIT researchers developed HardFlow, a deployment-time algorithm that lets pretrained generative models satisfy hard safety, physical, or task-specific constraints in their final output while still producing high-quality solutions, without retraining. The research appears in the IEEE Transactions on Pattern Analysis and Machine Intelligence.

The Federal Laboratory Consortium selected AI-GUIDE, a portable ultrasound-based medical device developed by MIT Lincoln Laboratory and Massachusetts General Hospital, for its 2026 Excellence in Technology Transfer Award. The prototype is being transferred to the startup AutonomUS Medical Technologies, Inc.

Amazon Science research finds that LLM-based ML research agents largely avoid overfitting reused benchmark validation sets because successful strategies are highly compressible. Across eight datasets, 32-token prompts let fresh agents reproduce explorer performance, while deliberately overfit strategies failed the compression test.

NVIDIA and Palantir built a Digital Supply Chain Intelligence command center on Palantir Foundry, using NVIDIA cuOpt for weekly allocation optimization and post-training a 30B Nemotron 3.5 Lightning model on captured planner decisions. On a development benchmark, the post-trained model reached 86.7% allocation-decision accuracy.

MIT Schwarzman College of Computing hosted its inaugural AI Educators Pilot, a weeklong July workshop for 19 faculty from seven universities, based on the Modeling with Machine Learning course. Participants explored adapting the course's materials and teaching methods for their own disciplines.

Google DeepMind and Primordial Soup collaborated on "Love, Rendered," a documentary short that uses AI to recreate an unrecorded memory for Burt and Ethelle Shatz, married over 70 years. The film was directed by Liz Garbus and produced by Dan Cogan and Darren Aronofsky.

Google DeepMind announced the AlphaGenome Atlas on 8 September, an online repository of precomputed AlphaGenome predictions covering all 9 billion possible single-letter changes to a reference human genome. It is freely available for noncommercial research, with potential commercial licensing.

A Hugging Face blog post describes an open reproduction of a project that trains a coding model to paint watercolours by writing p5.brush JavaScript, using TRL and OpenEnv with three reward mixes. The blog reports trained models and datasets published openly on the Hub.

A MIT News feature profiles three IBM researchers—Srinivasan Arunachalam, Zhang-Wei Hong PhD ’25, and Irene Ko PhD ’24—who credit the MIT-IBM Computing Research Lab with helping translate theoretical work into industry applications in quantum machine learning, reinforcement learning, and trustworthy AI.

MIT and Motional researchers developed the Concept-Wrapper Network (CW-Net), a method that translates a self-driving car's deep learning decisions into understandable concepts without altering driving performance. In private-track tests and simulation studies, CW-Net explanations helped safety drivers and nonexpert users better predict vehicle behavior. The research appears in Nature.

Researchers introduced BenchMIRT, a multidimensional Item Response Theory method for auditing LLM benchmarks at the level of individual prompts. Trained on results from 100 LLMs across 16 benchmarks and more than 34K questions, it independently recovered safety and general reasoning as dominant dimensions.