Introducing Falcon ASR

The Technology Innovation Institute (TII) in Abu Dhabi introduced Falcon-ASR, a 1.6 billion parameter speech recognition model built for Arabic with particular attention to the Emirati dialect. It also handles English, French, Spanish and Portuguese, and is available to try through a Hugging Face Demo Space, with API access and native applications planned.
On the Open Universal Arabic ASR Leaderboard protocol, which averages word error rate equally across six Arabic test sets, Falcon-ASR scored 20.92% WER. The best published result in the leaderboard snapshot the team checked on 30 September 2026 was 23.17%, putting Falcon-ASR 2.25 percentage points ahead. Word error rate measures the share of words transcribed incorrectly, and character error rate does the same at the character level; lower is better for both. On TII's internal Emirati evaluation, the model reached 22.73% WER and 10.19% CER, described as the lowest of the compared systems and 4.07 percentage points below Qwen3-Omni, the next best.
The model was trained on Emirati, Modern Standard Arabic, other Gulf and Arabic dialects, and English. Training data covered background noise, overlapping speech, music, room reverberation and telephony effects, plus speed and pitch variation, with the same treatment applied to Emirati recordings. Falcon-ASR also produces word-level timestamps linking each transcribed word to its position in the audio. For English, it recorded a mean WER of 5.74% across the seven public test sets used by the Hugging Face Open ASR Leaderboard. All five languages share the same weights and need no language flag; the output is a transcript in the spoken language.
The model builds on TII's Falcon3-Audio work, whose architecture and single-stage training approach are described in a paper on competitive audio-language models. Falcon-ASR's public Arabic evaluation used the leaderboard's pinned manifests, while the Emirati assessment relied on held-out recordings with human-validated transcripts.
Why it matters: Arabic speech recognition has had fewer transcribed resources for dialects than for Modern Standard Arabic, and this release targets that gap with a single set of weights covering five languages. Teams working with Arabic audio, including Emirati and Gulf speech, can test the model now through the demo, though API access and native applications are only planned.
Based on reporting from the original publisher. Visit the source for full context and later updates.