AivexaNewsSearch
AI news for builders and product teamsChecked every hour
Hugging Face BlogFirst partyResearch

Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

Collected Sep 30, 2026

Hugging Face has launched the Open TTS Leaderboard, an evaluation effort focused on open-source and multilingual text-to-speech and voice cloning models. The company said the Hub hosted more than 8K TTS models as of Sep 30, 2026, while evaluation has remained fragmented and unstandardized.

The leaderboard uses objective metrics rather than human votes: intelligibility via word and character error rate between the prompt and the transcript of generated audio, using Qwen3 ASR; speed via inverse real-time factor for batched offline inference on an H200 GPU and time-to-first-audio for streaming at batch size 1 on H200 GPU and CPU; and speaker similarity via cosine similarity between WavLM speaker embeddings of generated audio and the reference clip. Hugging Face said this cuts evaluation time from about two weeks of vote collection to a couple of hours.

Hugging Face stated the leaderboard does not replace human preference ranking. ASR-based WER is a proxy for intelligibility and speaker similarity estimates voice identity preservation; neither directly measures naturalness, expressiveness or listener preference. It said the metrics can inform voting-based leaderboards about which models to include.

The default view ranks models by macro-average WER on English splits of Seed TTS Eval and CV3 Eval (zero shot). hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro lead on English WER in that average. Seed TTS Eval only has English and Chinese audio, so other languages use CV3 Eval (zero shot) scores; Chinese, Japanese and Korean are character-based and reported as CER. k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 are described as strong multilingual models.

A Voice cloning toggle lets users compare models supporting it on selected languages, adds a SIM column and two Pareto plots for SIM, batched inference and size. Average WER for some models, including bosonai/higgs-tts-3-4b and openbmb/VoxCPM2, improves under voice cloning, when reference audio is provided. Pareto plots also visualize WER, RTFx and size.

A Listen tab lets users compare generated outputs behind the metrics, choose language/dataset, voice cloning and models, or sample randomly. Users can vote after logging in with an HF account to limit spam and bots; Hugging Face said vote data may appear on the leaderboard as more are collected.

A Streaming tab ranks models by time-to-first-audio, measured on the same 50 English CV3-Eval prompts, batch size 1, default voice and identical hardware, dropping the first three runs as warm-up and reporting the median. Default results are for an H200 GPU, with CPU results for a small but growing set. kyutai/pocket-tts performs well for streaming on GPU and CPU. Hugging Face said it will soon open-source evaluation scripts and seeks feedback via GitHub Issues and PRs.

Read at Hugging Face Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt