AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance

Collected Oct 6, 2026

The Technology Innovation Institute introduced Falcon-Emirati-7B, a model specialized for Emirati Arabic built on its Falcon-H1-Arabic family, according to a Hugging Face blog post. The model targets the vocabulary, tone and cultural context of the Gulf dialect rather than Modern Standard Arabic alone.

The base family uses a Falcon-H1 hybrid architecture combining State Space Models (Mamba) and Transformer attention in parallel within each block, with outputs fused before each block's projection. The family spans 3B, 7B and 34B parameters with context windows up to 128K and 256K tokens and was trained on MSA and dialectal Arabic (Gulf, Levantine, Egyptian, Maghrebi) alongside English and multilingual data.

The post states the 7B variant was chosen as a balance between quality and training and inference cost. Data came from three sources: crawled Emirati-dialect web content; MSA material about Emirati culture, heritage and identity; and synthetic dialect data constrained by Emirati glossaries and style rules. The institute said it ran ablations on data volume, training stage, authentic-versus-synthetic balance, and MSA cultural context, using automatic scoring and native-speaker review.

Evaluation used native-speaker manual review and Alyah, a 1,173-sample native multiple-choice benchmark for Emirati-dialect capability in Arabic LLMs. Falcon-Emirati-7B scored 84.83% on Alyah, ahead of every other Arabic and multilingual model compared, with the Falcon-H1-Arabic family excluded from that comparison.

On open-ended generation over the same 1,173 questions judged by Gemini 3.7 Flash against ALLaM-7B-Instruct-preview, gemma-3-27b-it, Jais-2-8B-Chat and Fanar-2-27B-Instruct, the post reports dialect fidelity of 0.52 partial credit for Falcon-Emirati-7B, versus 0.05 for ALLaM, 0.03 for gemma-3-27b-it, 0.02 for Jais-2-8B-Chat and effectively 0.00 for Fanar-2-27B-Instruct. Fanar-2-27B-Instruct abstained on 26.2% of questions, versus under 5% for every other model, with a correctness score of 0.27 partial credit, the lowest of the five.

Pairwise judging against Jais-2-8B-Chat, ALLaM-7B-Instruct-preview and Fanar-2-27B-Instruct showed Falcon-Emirati-7B winning most categories, with the widest margins in Poetry & Creative Expression and Language & Dialect. Greetings & Daily Expressions was the one category where competitors held their own.

On the UAE portion of ArabCulture-Dialogue, using 283 scenarios in Emirati Arabic and MSA across location settings, Falcon-Emirati-7B scored 85.57%, ahead of ALLaM-7B (83.39%), Jais-2-8B (73.79%) and Fanar-2-27B (71.50%).

The post notes the model can reflect training-data biases and err on rare expressions or highly localized references, and recommends evaluating it before sensitive, official or high-stakes use. Falcon-Emirati-7B is available on the Falcon chat platform. The authors are Shaikha Alsuwaidi, Omar Alkaabi, Maitha Alhammadi, Hamza Alobeidli, Ahmed Alzubaidi, Mohammed Alyafeai, Leen AlQadi, Basma Boussaha and Hakim Hacid, with 2026 given as the citation year.

Read at Hugging Face Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt