AI news for builders and product teamsUpdated Oct 10, 2026, 22:01 UTC
Research news
The latest Research stories across our sources, prepared from the publishers’ own reporting.
In this topic
Newest first
Nathan Lambert, author of a new post-training textbook on reinforcement learning from human feedback, argues that LLMs have stagnated at long-form non-fiction writing while progressing rapidly in coding and math. He used models for editing, LaTeX, and diagrams, but says only 10-20% of effort can be saved today.

A large-scale trial in Pakistan found that a custom AI tool for judges, combined with training, increased cases resolved by 6.3 percent with no obvious drop in judgment quality. The tool, JudgeGPT, combined OpenAI's GPT-4 with a knowledge base of Pakistani judicial opinions and statutes.

Amazon's Automated Reasoning Group reflects on a decade since its 2016 launch, describing how formal verification tools such as Tiros and Zelkova became AWS services including IAM Access Analyzer, Amazon Inspector, and Reachability Analyzer. The group says its production services process billions of queries daily and that it proved the correctness of infrastructure including the Nitro Isolation Engine and an authorization engine handling one billion API calls per second.

Researchers at MIT CSAIL and Tsinghua University developed GeoPT, a pre-training approach that helps simulation models learn physics from synthetic dynamics data. The team says it reaches peak performance twice as fast and trains on up to 60 percent less data than leading models.

Amazon announced 34 recipients of its Build on Trainium awards, part of a $110 million credit program supporting AI research at 30 universities. The Fall 2025 cycle focused on Responsible AI across five priority topics, all leveraging AWS Trainium infrastructure.

MIT researchers led by Ju Li published a Joule paper showing that a sulfonamide-family solvent called DMFSA can improve sodium-metal battery stability and fast cycling. An AI-guided algorithm designed 100,000 candidate molecules; 27 were tested, and DMFSA was the smallest and best.

A new study by MIT and other researchers found that AI assistance improved skin disease diagnosis accuracy for non-experts and clinicians, but explainability methods had different impacts depending on user knowledge level. Non-experts deferred to AI, especially LLM explanations, while clinicians performed best with only a model prediction and no explanation.

Amazon researchers presented ControlG at ICML, a framework that uses PID controllers from industrial control systems to schedule multiple training objectives in graph self-supervised learning by allocating compute to one objective at a time instead of blending gradients. It outperformed baselines on nine graph benchmarks across node classification, link prediction, and clustering.

Epoch and METR released MirrorCode, a benchmark testing whether AI systems can reimplement programs from CLI access alone; some tasks were solved, but 8 of 25 targets were never fully solved. Anthropic, Sunday, and OpenAI also reported robotics and safety findings.

The UK AI Security Institute found the cybersecurity capability gap between open-weight and closed frontier models has narrowed, with GLM-5.2 and DeepSeek V4-Pro performing like closed models released four to seven months earlier. Kimi also announced Kimi K3, a 2.8 trillion parameter model, while Demis Hassabis proposed a FINRA-style standards body for frontier AI testing.

An article from Ahead of AI explains how large language models can be trained to support multiple reasoning-effort modes, citing OpenAI's GPT-5.6 family, gpt-oss, Qwen3, and Thinking Machine Labs' Inkling as examples.

Amazon and the University of Michigan developed HydroShear, a physics-based simulator that models tactile shear forces so robots can learn contact-rich manipulation policies in simulation that transfer to real hardware. It achieved a 93% average success rate across four tasks on a Franka robot with GelSight Mini sensors, versus 34% for TacSL and 58-61% for FOTS.

Import AI 464 covers an AI-written GPU megakernel on KernelBench-Mega, rising AI automation on the Remote Labor Index, the OSWorld 2.0 computer-use benchmark, and JD's Oxygen AI Item Center.

Lilian Weng published a post on harness engineering for recursive self-improvement, covering harness design patterns, context engineering methods such as ACE and MCE, automated workflow search approaches including ADAS and AFlow, and Meta-Harness. The post also traces the concept of recursive self-improvement to I. J. Good (1965) and Yudkowsky (2008).

Together AI says it will present nine papers at ICML 2026, spanning agent evaluation, model training, inference optimization, and GPU kernels, with some research shipping in its production platform. The company will be at booth B714 in Seoul from July 6 to 11.

NVIDIA researchers have developed ENPIRE, a framework that lets coding agents autonomously refine robot policies in the real world. Separately, Tencent detailed ARGUS, tracing software used on a production cluster of over 10,000 GPUs for more than six months.

A technical article by Lilian Weng reviews the history and methodology of neural scaling laws, tracing them from early learning-curve work through Kaplan et al. (2020), the Chinchilla scaling laws, and efforts to reconcile the two. It also covers scaling in data-limited regimes and practical difficulties in fitting scaling laws.

Together AI has released ParallelKernelBench, a benchmark testing whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves fewer than a third of problems, but a few generated kernels beat any public implementation.

A study across four experiments with 18,978 conversations found AI systems were reliably more persuasive than expert humans, including professional canvassers, and nearly 3x more effective at raising real-money donations to Save the Children. Constraining AI to human-length messages at human writing speeds eliminated its advantage over coached elite debaters.

Researchers built SocioHack, a 72-environment benchmark showing RL-trained models rediscover historically patched regulatory loopholes; Anthropic reports an 8x increase in code merged in 2026 versus 2021-2024; and RL-trained quadrotors beat a champion human pilot.