AivexaNewsSearch
AI news for builders and product teamsUpdated Oct 10, 2026, 22:01 UTC

Research news

The latest Research stories across our sources, prepared from the publishers’ own reporting.

In this topic

Newest first

Physics AI research that’s shaping the industry.

Mistral AI published a blog post outlining physics AI research areas it is pursuing following its acquisition of Emmi AI, listing several papers and models covering transonic aerodynamics, computational fluid dynamics, automotive and aerospace applications, plasma turbulence, and industrial simulation. Current scientific results from these efforts are sparse and remain at a preliminary stage.

Read original
OllamaFirst partyResearch

NVIDIA DGX Spark performance

Ollama published benchmark results for running various models on NVIDIA DGX Spark hardware using firmware 580.95.05 and Ollama v0.12.6, reporting prefill and decode tokens-per-second figures across models including gpt-oss, gemma3, llama3.1, deepseek-r1, and qwen3.

Read original

Real AI Agents and Real Work

OpenAI released a new expert-designed test measuring whether AI can perform economically relevant work tasks, and human experts won only narrowly, with AI failures mostly in formatting and instruction-following rather than hallucinations. The author also describes giving Claude Sonnet 4.5 an economics paper and its replication data, prompting it to attempt a replication.

Read original

AGI Is Not Multimodal

An essay in The Gradient argues that multimodal AI, which stitches together separate modality modules, will not reach human-level AGI in the near term, and that embodiment and environmental interaction should be treated as primary instead.

Why We Think

Lilian Weng published a review of test-time compute and chain-of-thought reasoning, covering parallel sampling, sequential revision, reinforcement learning, and continuous-space thinking. She credits John Schulman with feedback and edits.

Agents

Chip Huyen published a blog post adapted from the Agents section of her book AI Engineering (2025). It covers how agents are defined, the role of tools in extending their capabilities, and the challenges of planning for complex tasks.

Reward Hacking in Reinforcement Learning

In a new post, Lilian Weng surveys reward hacking in reinforcement learning, where agents exploit flaws or ambiguities in reward functions to score highly without completing the intended task. She catalogues examples across RL, language model and real-world settings, and calls for more research on practical mitigations, especially for RLHF and LLMs.

Read original

Shape, Symmetries, and Structure: The Changing Role of Mathematics in Machine Learning Research

An article in The Gradient argues that mathematics remains central to machine learning research even as large-scale, engineering-first methods have produced breakthroughs that outpace existing theory. It says mathematics' role is evolving toward post-hoc explanation and higher-level design guidance, with tools such as intrinsic dimension, curvature and topology now applied to deep learning models.

Read original

Extrinsic Hallucinations in LLMs

Lilian Weng published an article on extrinsic hallucinations in large language models, defining them as fabricated outputs not grounded in pre-training data or world knowledge. The post covers causes, detection methods, and anti-hallucination techniques, citing research including Gekhman et al. 2024, FactualityPrompt, FActScore, SAFE, SelfCheckGPT, TruthfulQA, and SelfAware.

Diffusion Models for Video Generation

Lilian Weng's blog post explains how diffusion models are being extended from image synthesis to video generation. It covers training video diffusion models from scratch, including parameterization, sampling, 3D U-Net and DiT architectures, and adapting pre-trained image models to video.

Read original

A Brief Overview of Gender Bias in AI

The Gradient article surveys research on measuring gender bias in AI models, covering word embeddings, facial recognition, coreference resolution, question answering, and image generation. It also discusses gaps in the research and open questions about how, or whether, such bias should be fixed.