AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

Collected Oct 1, 2026

Apple Machine Learning Research published a paper titled "Compressing Streaming Neural Audio Encoders via Latent-Space Distillation." The authors are Prasanth Yadla, Mohammad Samragh Razlighi, Dongseong Hwang, Mingbin Xu, Yuanyuan Zhang, Chung-Cheng Chiu, Yongqiang Wang, Yuan Liu, Zhen Huang and Xiaodan Zhuang. Yadla and Razlighi are listed as equal contributors; Wang is affiliated with NVIDIA, Liu with Anthropic, and Liu's work was done while at Apple.

System-wide Dictation on Apple devices runs entirely on-device, according to the paper. Speech reaches the foundation model through a tokenizer, an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency.

The work studies compressing such a tokenizer by distillation, using as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes, described as the last representation the two token interfaces share. Only the student encoder is trained, regressing the teacher's per-frame latent under a squared-error objective, with a single affine layer absorbing the teacher-student width mismatch. Because the target precedes both the quantizer and the language-model bridge, one recipe covers both supported token interfaces and applies both to a tokenizer pretrained alone and to one jointly trained with a language model.

At 2.8x compression, the distilled student stays within 1.9% relative WER of its teacher on five of six teacher-student pairs without any fine-tuning, and improves on an independently trained tokenizer of identical capacity by 3.9% relative. The paper is dated September 2026.

Read at Apple Machine Learning Research

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study