Compressing Streaming Neural Audio Encoders via Latent-Space Distillation
Apple Machine Learning Research published a paper titled "Compressing Streaming Neural Audio Encoders via Latent-Space Distillation." The authors are Prasanth Yadla, Mohammad Samragh Razlighi, Dongseong Hwang, Mingbin Xu, Yuanyuan Zhang, Chung-Cheng Chiu, Yongqiang Wang, Yuan Liu, Zhen Huang and Xiaodan Zhuang. Yadla and Razlighi are listed as equal contributors; Wang is affiliated with NVIDIA, Liu with Anthropic, and Liu's work was done while at Apple.
System-wide Dictation on Apple devices runs entirely on-device, according to the paper. Speech reaches the foundation model through a tokenizer, an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency.
The work studies compressing such a tokenizer by distillation, using as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes, described as the last representation the two token interfaces share. Only the student encoder is trained, regressing the teacher's per-frame latent under a squared-error objective, with a single affine layer absorbing the teacher-student width mismatch. Because the target precedes both the quantizer and the language-model bridge, one recipe covers both supported token interfaces and applies both to a tokenizer pretrained alone and to one jointly trained with a language model.
At 2.8x compression, the distilled student stays within 1.9% relative WER of its teacher on five of six teacher-student pairs without any fine-tuning, and improves on an independently trained tokenizer of identical capacity by 3.9% relative. The paper is dated September 2026.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study