TII Releases Falcon-Emirati-7B: A Dialect-Specialized LLM for Emirati Arabic
Key point
The new 7B model achieves 84.83% on the Alyah benchmark and significantly outperforms larger general-purpose models in dialect fidelity and cultural understanding.
Details
Technology Innovation Institute (TII) has released Falcon-Emirati-7B, a specialized Large Language Model designed to understand and generate the Emirati Arabic dialect. Built on the Falcon-H1-Arabic hybrid architecture, the model combines State Space Models (Mamba) for linear time efficiency with Transformer attention for long-range dependency precision. The 7B parameter size was selected as the optimal balance between capturing cultural nuance and maintaining practical training and inference costs.
Dialect Adaptation and Data Strategy
Standard Arabic (MSA) models often fail to capture the idioms, humor, and cultural context of spoken Emirati Arabic. To address this, TII developed a three-part data pipeline:
- Authentic Emirati-Dialect Web Data: Native text crawled from UAE websites and forums.
- MSA Data About Emirati Culture: Texts covering UAE history, values, and social norms to provide thematic knowledge.
- Synthetic Data: Generated using strict lexicons and style rules to prevent inaccurate "Gulf-ish" Arabic.
Benchmark Performance
The model was evaluated against several leading Arabic and multilingual models, including ALLaM-7B-Instruct-preview, gemma-3-27b-it, Jais-2-8B-Chat, and Fanar-2-27B-Instruct.
- Alyah Benchmark: Falcon-Emirati-7B scored 84.83%, surpassing all compared models, including those with significantly larger parameter counts.
- Dialect Fidelity: In open-ended generation tests judged by Gemini 3.7 Flash, Falcon-Emirati-7B achieved a dialect fidelity score of 0.52, compared to 0.05 for ALLaM, 0.03 for gemma-3-27b-it, 0.02 for Jais-2-8B-Chat, and 0.00 for Fanar-2-27B-Instruct. Competitors tended to default to MSA even when answering correctly.
- Cultural Understanding: On the ArabCulture-Dialogue UAE benchmark, Falcon-Emirati-7B achieved 85.57% accuracy, leading ALLaM-7B (83.39%), Jais-2-8B (73.79%), and Fanar-2-27B (71.50%).
Key Insights
The results demonstrate that model size alone does not improve dialect capability. Specialized training on authentic dialect data and cultural context is critical for handling nuances like Nabati poetry, proverbs, and social etiquette. While the model excels in dialect fidelity, TII notes limitations in rare expressions and highly localized references, recommending evaluation for sensitive or official use cases.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.