AivexaNewsSearch
AI news for builders and product teamsUpdated Oct 10, 2026, 19:01 UTC

Research news

The latest Research stories across our sources, prepared from the publishers’ own reporting.

In this topic

Newest first

The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models

Apple researchers propose a round-trip protocol to measure how much tree-structured compositional content survives when language models serialize arithmetic expressions into natural language. Testing all pairwise combinations of sixteen models, they report lossy and asymmetric communication, with generation as the dominant failure source, and show the channel is trainable with about 3600 fine-tuning examples.

Read original

Estimating suicide risk from text

MIT McGovern Institute researchers developed a language-processing tool that uses a custom lexicon tied to 49 suicide risk factors to estimate suicide risk from text conversations with crisis counselors. Reported in the Journal of Psychopathology and Clinical Science, it accurately predicted risk in about 16,000 de-identified Crisis Text Line conversations and may aid assessment with more validation.

Read original

A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization

Apple researchers present a recipe for semi-supervised federated ASR that pairs online pseudo-labels from a per-client teacher with server-side updates on labeled data to stabilize training. The method improves over the strongest prior approach on 9 of 11 pairs, by 20.8% on average in-domain and 10.0% cross-domain.

Read original

Can Jev Be a Better Agent Evaluator?

LangChain tested TypeSafe AI's Jev as an agent evaluator against GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6, finding Jev matched the human oracle on all 500 binary decisions, had 92–913x lower quality-score variance, and cost $0.00035 per call. LangChain calls the results promising but early.

Read original

[AINews] Here are 6 Clones of Jev in 2 days

Six reproductions and clones of the non-open-source Jev decision model appeared within two days of its launch, which drew 36M views on its launch video. The clones use methods including ModernBERT encoders, diffusion models, LoRA fine-tunes of Qwen3.5-9B and Qwen2.5-0.5B, and a 4B/35B Qwen3.5 backbone with an NLI classifier.

Read original

Dynamically Scaled Activation Steering

Apple researchers introduced Dynamically Scaled Activation Steering (DSAS), a method-agnostic framework that adaptively modulates the strength of existing activation steering transformations across layers and inputs. When combined with existing steering methods, DSAS reportedly improves the trade-off between toxicity mitigation and utility preservation and adds minimal computational overhead.

Read original