AI news for builders and product teamsUpdated Oct 10, 2026, 18:01 UTC
Apple Machine Learning Research
First-party releases and research from Apple Machine Learning Research. Headlines and excerpts link to the original articles.
Latest stories
Newest first
Apple researchers introduce Normalizing Trajectory Models (NTM), which model each reverse diffusion step as a conditional normalizing flow with exact likelihood training. On text-to-image benchmarks, NTM matches or outperforms strong baselines in just four sampling steps.

Apple researchers present RISED, a method that uses rubrics—textual descriptions of rollout behaviour—to guide data selection and policy supervision for training a single LLM agent across diverse interactive environments. RISED reportedly achieves the highest mean pass rate across environments and ranks first or second in each individual environment.

Apple researchers and Stanford collaborators designed two open-ended probes using a Wizard of Oz technique to let people train personalized machine learning systems on phenomena they define themselves. A week-long study identified four sites where ontological boundaries were negotiated.

Apple researchers show that discrete diffusion samplers match the training distribution only when the positions written per step are conditionally independent given already-fixed tokens, and that no product of per-position distributions can match a dependent group.

Apple researchers report that strengthening language discrimination during pretraining narrows the multilingual gap in English/French HuBERT speech models, improving phone discrimination, lexical and prosodic measures. Gains were largest when the intervention was introduced in the first training iteration.

An Apple Machine Learning Research paper reports that, under an equal time budget and the same frontier LLM backbone, open-source state-of-the-art harnesses provided no advantages over a single session of a minimal-harness coding agent baseline. The authors argue the backbone is the primary driver of performance.

Apple researchers present RLTL;DR, a method for self-improvement in reinforcement learning when tasks are so hard the agent rarely succeeds and no teacher models exist. It has the policy write its own TL;DR feedback after failed attempts, reaching 12-13% Pass@1 at evaluation on tool-calling and coding datasets.

Apple Machine Learning Research published a paper describing SCLATE, an execution substrate that lets benchmarks and unmodified agents add events to one open event scheduler through an adapter. The paper reports porting seven benchmarks, comparing ten unmodified harness and memory configurations on ten models, and post-training Qwen3.5-4B through unmodified harnesses and memory systems.

An Apple Machine Learning Research paper accepted at EMNLP systematically studies conditioning methods for controlling LLM outputs, finding that efficient steering methods often condition effectively at a steep cost to fluency. It also reports that activation steering is far less effective on instruction-tuned models than on base models.

Apple researchers propose a round-trip protocol to measure how much tree-structured compositional content survives when language models serialize arithmetic expressions into natural language. Testing all pairwise combinations of sixteen models, they report lossy and asymmetric communication, with generation as the dominant failure source, and show the channel is trainable with about 3600 fine-tuning examples.

Apple researchers present improved convergence rates for federated optimization of stochastic variational inequalities, including a refined analysis of Local Extra SGD and a new algorithm, LIPPAX, aimed at reducing client drift.

Apple researchers present a recipe for semi-supervised federated ASR that pairs online pseudo-labels from a per-client teacher with server-side updates on labeled data to stabilize training. The method improves over the strongest prior approach on 9 of 11 pairs, by 20.8% on average in-domain and 10.0% cross-domain.

Apple researchers distilled streaming neural audio encoders by matching the teacher's pre-quantizer latent rather than tokens or output distributions. At 2.8x compression, the student stayed within 1.9% relative WER of the teacher on five of six pairs without fine-tuning.

Apple researchers introduce probe guidance, a method that uses frozen internal states of an existing diffusion model to construct a guidance signal for flow matching models without an extra forward pass at inference. It sets a new state-of-the-art on unconditional generation for continuous diffusion language models.

Apple researchers introduced Dynamically Scaled Activation Steering (DSAS), a method-agnostic framework that adaptively modulates the strength of existing activation steering transformations across layers and inputs. When combined with existing steering methods, DSAS reportedly improves the trade-off between toxicity mitigation and utility preservation and adds minimal computational overhead.