AI news for builders and product teamsUpdated Oct 10, 2026, 22:01 UTC
Research news
The latest Research stories across our sources, prepared from the publishers’ own reporting.
In this topic
Newest first
Import AI 459 covers a paper estimating the US AI economy's quality-adjusted output grew roughly 2,290 percent in 2024 and 2,271 percent in 2025, UK AISI research on why automated alignment is difficult, Stanford-led release of the 100M-image GPIC dataset, and Biohub's ESMFold2 protein model.

Mistral AI published a blog post outlining physics AI research areas it is pursuing following its acquisition of Emmi AI, listing several papers and models covering transonic aerodynamics, computational fluid dynamics, automotive and aerospace applications, plasma turbulence, and industrial simulation. Current scientific results from these efforts are sparse and remain at a preliminary stage.

Import AI 457 covers SentinelOne's teardown of the fast16.sys virus that tampered with high-precision calculation software, Tilde Research's finding that the Muon optimizer can cause neuron death in MLP layers and its proposed Aurora optimizer, a position paper on positive alignment, and Prime Intellect tests of autonomous AI research agents on the nanoGPT speedrun.

A technical article surveys recent open-weight LLM architecture changes aimed at long-context efficiency from April to May, including KV sharing and per-layer embeddings in Gemma 4, attention budgeting in Laguna XS.2, compressed convolutional attention in ZAYA1-8B, and mHC plus compressed attention in DeepSeek V4.

Researchers with the Institute for Law & AI propose "radical optionality" for AI governance, urging governments to build institutions and legal authorities now while avoiding overregulation. Separately, a Meta and KAIST paper explores neural computers, and economists model how automating AI research could produce explosive economic growth.

Import AI editor Jack Clark writes that there is a 60%+ chance that no-human-involved AI R&D, where a system could autonomously build its own successor, happens by the end of 2028, citing benchmark trends in coding, reproducibility, ML engineering, and kernel design.

The essay revisits the "Jagged Frontier" concept of uneven AI ability, arguing bottlenecks can block automation while reverse salients like Google's Nano Banana Pro image model can suddenly remove them.

Ollama published benchmark results for running various models on NVIDIA DGX Spark hardware using firmware 580.95.05 and Ollama v0.12.6, reporting prefill and decode tokens-per-second figures across models including gpt-oss, gemma3, llama3.1, deepseek-r1, and qwen3.

OpenAI released a new expert-designed test measuring whether AI can perform economically relevant work tasks, and human experts won only narrowly, with AI failures mostly in formatting and instruction-following rather than hallucinations. The author also describes giving Claude Sonnet 4.5 an economics paper and its replication data, prompting it to attempt a replication.

An essay in The Gradient argues that multimodal AI, which stitches together separate modality modules, will not reach human-level AGI in the near term, and that embodiment and environmental interaction should be treated as primary instead.

Lilian Weng published a review of test-time compute and chain-of-thought reasoning, covering parallel sampling, sequential revision, reinforcement learning, and continuous-space thinking. She credits John Schulman with feedback and edits.

Chip Huyen published a blog post adapted from the Agents section of her book AI Engineering (2025). It covers how agents are defined, the role of tools in extending their capabilities, and the challenges of planning for complex tasks.

In a new post, Lilian Weng surveys reward hacking in reinforcement learning, where agents exploit flaws or ambiguities in reward functions to score highly without completing the intended task. She catalogues examples across RL, language model and real-world settings, and calls for more research on practical mitigations, especially for RLHF and LLMs.

An article in The Gradient argues that mathematics remains central to machine learning research even as large-scale, engineering-first methods have produced breakthroughs that outpace existing theory. It says mathematics' role is evolving toward post-hoc explanation and higher-level design guidance, with tools such as intrinsic dimension, curvature and topology now applied to deep learning models.

A Gradient essay argues that beneficial AI should be grounded in human wellbeing, that positive visions for an AI-infused society are needed, and that foundation models and their future deployment are critical leverage points.

Lilian Weng published an article on extrinsic hallucinations in large language models, defining them as fabricated outputs not grounded in pre-training data or world knowledge. The post covers causes, detection methods, and anti-hallucination techniques, citing research including Gekhman et al. 2024, FactualityPrompt, FActScore, SAFE, SelfCheckGPT, TruthfulQA, and SelfAware.

An article by Richard Dewey and Ciamac Moallemi in The Gradient examines whether large language models can be applied to financial market prediction, weighing obstacles like scarce data and market efficiency against potential uses such as multimodal learning and synthetic data.

Lilian Weng's blog post explains how diffusion models are being extended from image synthesis to video generation. It covers training video diffusion models from scratch, including parameterization, sampling, 3D U-Net and DiT architectures, and adapting pre-trained image models to video.

The Gradient article surveys research on measuring gender bias in AI models, covering word embeddings, facial recognition, coreference resolution, question answering, and image generation. It also discusses gaps in the research and open questions about how, or whether, such bias should be fixed.

Chip Huyen published an analysis of 845 open source AI software repositories with at least 500 GitHub stars, examining the stack around foundation models. She reports 2023 saw the largest growth in application and application development layers, and that China's open source AI ecosystem has diverged from the Western one.