AivexaNewsSearch
AI news for builders and product teamsChecked every hour

A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization

Collected Oct 1, 2026

Apple Machine Learning Research published a paper titled "A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization," listed under Methods and Algorithms and Speech and Natural Language Processing, with a publication date of September 2026. Authors are Wonho Bae, Zakaria Aldeneh, Martin Pelikan, Jan "Honza" Silovsky, Tatiana Likhomanenko and Sheikh Shams Azam.

The paper addresses semi-supervised federated learning (SSFL), in which models are trained on clients' unlabeled data using a teacher to generate pseudo-labels, alongside a small labeled seed dataset on the server. It states that Automatic Speech Recognition (ASR) is particularly fragile in this setting because pseudo-label errors compound across the output sequence and across training rounds into divergence, leaving a large gap to fully-supervised federated learning.

The work frames closing that gap around two coupled design axes: the teacher, meaning which model generates the pseudo-labels, and the anchor, the server-side updates on labeled data that stabilize training. On the teacher axis, a per-client online teacher (each client's own evolving model) diverges on its own, but once stabilized it matches or beats the broadcast global teacher, described as one server model fixed within a round—decisively in-domain and competitively under domain shift. As the seed grows stronger and the online teacher's advantage narrows, a transitioning teacher (global to online at round r) matches or beats both.

On the anchor axis, the server must keep training on labeled data between rounds, otherwise the online teacher drifts, and the paper reports that this interleaving, more than the seed model, governs convergence. It states the two axes are inseparable: aggressive teacher choices pay off only once the anchor stabilizes training, which is described as highly sensitive to data augmentation and batch size—the settings governing how much input and gradient noise the server injects. How much stabilization is needed is domain-dependent, governed by the dispersion of the seed data and its overlap with client data.

The paper reports that these findings yield guidelines for SSFL in ASR training, improving over the strongest prior method on 9 of 11 pairs, by 20.8% on average in-domain and 10.0% cross-domain, narrowing the gap to fully-supervised federated learning.

Read at Apple Machine Learning Research

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Semi-supervised federated learning (SSFL) trains models on clients’ unlabeled data using a teacher to generate pseudo-labels, with a small labeled seed dataset on the server. Automatic Speech Recognition (ASR) is particularly fragile here: pseudo-label errors compound across the output sequence and across training rounds into divergence, leaving a large gap to fully-supervised FL. We show that closing this gap turns on two coupled design axes—the teacher (which model generates the pseudo-labels) and the anchor (the