AivexaNewsSearch
AI news for builders and product teamsUpdated Oct 10, 2026, 23:01 UTC

Research news

The latest Research stories across our sources, prepared from the publishers’ own reporting.

In this topic

Newest first

Why Doesn’t My Model Work?

A Gradient article by an author with about 20 years of machine learning experience describes common ways models appear to perform well while failing on real-world data, including misleading data, data leakage and inappropriate metrics. It cites a review by Roberts et al. of Covid prediction models and a Toronto water quality system, and points to the REFORMS checklist.

Read original

Adversarial Attacks on LLMs

Lilian Weng published an overview post on adversarial attacks against large language models, covering threat models, attack classification, and specific techniques such as token manipulation, gradient-based attacks, and jailbreak prompting.

Read original

Neural algorithmic reasoning

A Gradient article surveys neural algorithmic reasoning, the effort to capture classical computation such as shortest path-finding in deep neural networks. It covers algorithmic alignment theory, graph neural networks for algorithm execution, and applications including mathematics and computer networking.

Read original

Multimodality and Large Multimodal Models (LMMs)

Chip Huyen published a long technical post explaining multimodality and large multimodal models (LMMs), covering why multimodal data matters, the different data modalities and multimodal tasks, and the fundamentals of systems like CLIP and Flamingo. It also surveys active LMM research areas including multimodal outputs and efficient training adapters.

Read original

Open challenges in LLM research

Chip Huyen outlines ten open research directions in large language model work, based on conversations with people in industry and academia. She highlights hallucinations and context learning as the most discussed areas, and names multimodality, new architectures and GPU alternatives as the ones she is most excited about.

Read original

Prompt Engineering

Lilian Weng published a post explaining prompt engineering, also called in-context prompting, as methods to steer LLM behavior without updating model weights. The post covers zero-shot and few-shot prompting, example selection and ordering, instruction prompting, self-consistency sampling, chain-of-thought prompting, and automatic prompt design.

Read original

The Transformer Family Version 2.0

Lilian Weng published version 2.0 of her "The Transformer Family" post, a refactored and expanded update to her 2020 article. The new version restructures the section hierarchy, adds more recent papers, and is described as a superset of the old version at about twice its length.

Read original

Large Transformer Model Inference Optimization

A technical blog post by Lilian Weng provides an overview of methods for optimizing inference of large transformer models, covering distillation, quantization, pruning, sparsity, mixture-of-experts, and architecture-specific improvements. It explains challenges such as high memory usage, including KV cache memory, and reviews techniques to reduce memory footprint, computation, and latency.

Read original

Some Math behind Neural Tangent Kernel

Lilian Weng published a technical blog post titled "Some Math behind Neural Tangent Kernel," providing a deep dive into the motivation, definition, and proofs behind neural tangent kernel (NTK) theory. The post covers NTK's connection to Gaussian processes and the proof of deterministic convergence for infinite-width networks.

Generalized Visual Language Models

Lilian Weng's post surveys one approach to vision language models: extending pre-trained language models to consume visual signals. It groups methods into four buckets and describes techniques such as jointly training image and text, learned image embeddings as frozen LM prefixes, cross-attention fusion, and training-free decoding guided by vision-based scores.

What are Diffusion Models?

Lilian Weng's blog post "What are Diffusion Models?" explains diffusion-based generative models, covering forward and reverse diffusion processes, connections to score networks and Langevin dynamics, guidance methods, and later updates adding latent diffusion, progressive distillation, and consistency models.

Read original