Oct 6 edition/Reporting & analysis
ModelsResearchBusiness

ModelsArchitectures & capability

TII's Falcon-Emirati-7B targets Emirati Arabic, but its edge rests on TII's own tests

The Technology Innovation Institute's Falcon-Emirati-7B, a fine-tune of Falcon-H1-Arabic, replies in Emirati dialect far more often than rival Arabic models in TII's LLM-judged tests. Its multiple-choice gain over its own base model is only about 2.7 points.

Illustration from Hugging Face: TII's Falcon-Emirati-7B targets Emirati Arabic, but its edge rests on TII's own tests
Image: Hugging Face — Original article ↗
THE CORE IDEAS4 TAKEAWAYS
01

The headline result is dialect fidelity. Rival Arabic models usually answered Emirati-dialect questions in Modern Standard Arabic. On an LLM-judged partial-credit score, Falcon-Emirati-7B scored 0.52. ALLaM, gemma-3-27b-it, Jais-2-8B-Chat and Fanar-2-27B each scored 0.05 or lower. In pairwise judging it won strongly on poetry but narrowly lost on greetings to Jais. [1]

02

On Alyah, TII's multiple-choice Emirati benchmark, the model scores 84.83%. Its base model, Falcon-H1-Arabic-7B-Instruct, scored 82.18% when the benchmark launched in January. The launch comparison leaves the base model out. As a result, the gain from specialization looks modest on knowledge questions and large only on reply register. [1] [4]

03

The training recipe combines three kinds of data: native Emirati web and forum text, Modern Standard Arabic material about Emirati culture, and synthetic dialect text constrained by glossaries and grammar rules. The model is built on a hybrid Mamba-attention base. TII does not disclose the final data mix, token counts or chosen training stage. The reviewed sources confirm access only through Falcon Chat, with no confirmed weights or license. [1] [2] [6]

04

Independent studies confirm the underlying problem: language models perform worse on Arabic dialects than on Modern Standard Arabic. None of these studies evaluates Falcon-Emirati itself. [5] [8] [9]

WHY IT MATTERS

independent studies show LLMs handle Arabic dialects worse than MSA. Implication: UAE deployments should test reply register separately from accuracy, since strong Arabic scores need not mean dialect output.

Read the full assessment

Falcon-Emirati's advantage there is so far vendor-reported.

Executive brief

Most Arabic chatbots know the answer to a question asked in Emirati Arabic, but they reply in Modern Standard Arabic (MSA), the formal written variety. In an LLM-judged test reported by the Technology Innovation Institute (TII), its new Falcon-Emirati-7B kept the dialect with a partial-credit score of 0.52. The four comparison models scored 0.05 or lower (TII blog). The model is a 7B fine-tune of Falcon-H1-Arabic and scores 84.83% on Alyah, a multiple-choice Emirati benchmark. Its own base model scored 82.18% on Alyah in January, and that base model is left out of the launch comparison. All results are vendor-reported. No independent evaluation was found.

What changed and event timeline

  1. Falcon-H1-Arabic launches

    TII released hybrid Mamba-Transformer Arabic models at 3B, 7B and 34B. They were trained on about 300B tokens that include Gulf and other dialects.

  2. Alyah benchmark released

    TII published 1,173 native Emirati multiple-choice items and evaluated 54 models. Falcon-H1-Arabic-7B-Instruct led the instruction-tuned models at 82.18%.

  3. ArabCulture-Dialogue preprint

    The preprint presents a 13-country conversational benchmark for cultural reasoning. Models did worse on dialect than on MSA across all three of its tasks.

  4. Falcon-Emirati-7B announced

    The model was announced with Alyah, LLM-judge and cultural-reasoning results, and is available through Falcon Chat.

  5. Part of a three-model release

    TII also announced Falcon-ASR (1.6B, speech recognition) and Falcon-OCR-Arabic.

Capabilities and access

  • Model: Falcon-Emirati-7B, a dialect-specialized model built on Falcon-H1-Arabic-7B (TII blog).
  • Purpose: understanding and generating Emirati Arabic, including idioms, proverbs, nabati poetry, etiquette and heritage knowledge.
  • Access: the reviewed sources document only the Falcon Chat web interface (TII blog; Arabian Reseller). Neither confirms downloadable weights or a license.
Read the full section
  • Model: Falcon-Emirati-7B, a dialect-specialized model built on Falcon-H1-Arabic-7B (TII blog).
  • Purpose: understanding and generating Emirati Arabic, including idioms, proverbs, nabati poetry, etiquette and heritage knowledge.
  • Access: the reviewed sources document only the Falcon Chat web interface (TII blog; Arabian Reseller). Neither confirms downloadable weights or a license. The base family was released as open models (Falcon-H1-Arabic blog).

Technical analysis for researchers and developers

  • Architecture: inherited from the base model. The 7B base has a 256K-token context (Falcon-H1-Arabic blog).
  • Data: three sources (TII blog):
  • Emirati web and forum text written natively in dialect
Read the full section
  • Architecture: inherited from the base model. Mamba (state-space) and attention layers run in parallel in each block, and their outputs are concatenated before the block's output projection. The 7B base has a 256K-token context (Falcon-H1-Arabic blog).
  • Data: three sources (TII blog):
  • Emirati web and forum text written natively in dialect
  • MSA material about Emirati culture
  • Synthetic dialect text, constrained by Emirati glossaries and grammar rules
  • Recipe: TII ran ablations on data mix and training stage (continued pre-training, supervised fine-tuning, preference optimization). It does not publish the final mix, token counts or which stage it chose.
  • Evaluation: Alyah multiple-choice accuracy, native-speaker review, and an LLM judge ("Gemini 3.7 Flash") scoring open-ended answers for correctness and dialect fidelity, plus blind pairwise comparisons. Prompts, judge rubrics and the synthetic-data generator are not released, so the results cannot be reproduced.

Claims and evidence

All of the following are vendor-reported. No independent evaluation was found.

Read the full section

All of the following are vendor-reported. No independent evaluation was found.

  • Alyah: 84.83%, ahead of every model TII compared it against. Falcon-H1-Arabic models are excluded from that chart (TII blog). The base 7B-Instruct model scored 82.18% in January (Alyah blog), so the multiple-choice gain over the base is about 2.7 points.
  • Dialect fidelity (partial credit): 0.52, against 0.05 for ALLaM, 0.03 for gemma-3-27b-it, 0.02 for Jais-2-8B-Chat and about 0.00 for Fanar-2-27B (TII blog).
  • Pairwise judging: wins Poetry 0.88–0.12 against Fanar. Loses Greetings 0.46–0.54 to Jais.
  • ArabCulture-Dialogue, UAE subset (283 scenarios, TII's own run): 85.57%, against 83.39% for ALLaM-7B (TII blog). The benchmark itself comes from an outside group (arXiv).

Context and prior work

  • Falcon-H1-Arabic already claimed the top Open Arabic LLM Leaderboard (OALL) results, with 71.7% for the 7B model (Falcon-H1-Arabic blog; TII news).
  • Independent research finds that models handle Arabic dialects worse than MSA:
  • ArabCulture-Dialogue (arXiv)
Read the full section
  • Falcon-H1-Arabic already claimed the top Open Arabic LLM Leaderboard (OALL) results, with 71.7% for the 7B model (Falcon-H1-Arabic blog; TII news).
  • Independent research finds that models handle Arabic dialects worse than MSA:
  • ArabCulture-Dialogue (arXiv)
  • DialectalArabicMMLU (arXiv)
  • AL-QASIDA (arXiv)
  • These studies support the problem TII describes. None of them tests Falcon-Emirati.

Limitations, safety and contested findings

  • Evaluation independence: Alyah was built largely by the same TII authors who built the model (Alyah blog). TII also ran every comparison itself.
  • Judge validity: a single proprietary LLM judge scored the open-ended results, and no agreement with human raters is reported.
  • Omitted baseline: leaving out the base model makes the gain from specialization hard to see.
Read the full section
  • Evaluation independence: Alyah was built largely by the same TII authors who built the model (Alyah blog). TII also ran every comparison itself.
  • Judge validity: a single proprietary LLM judge scored the open-ended results, and no agreement with human raters is reported.
  • Omitted baseline: leaving out the base model makes the gain from specialization hard to see. On multiple choice it appears modest; the large gap is in dialect fidelity.
  • TII's own caveats: the model can reflect training-data biases, makes errors on rare or highly local expressions, and native speakers sometimes disagree on the right answer. TII advises testing before any sensitive or official use (TII blog).

Business and practitioner implications

  • Dialect is a separate requirement. Teams should measure the register of replies separately from accuracy.
  • TII's documented approach (native web text, culture-focused MSA, glossary-constrained synthetic data, a native-speaker benchmark) could be applied to other low-resource dialects at 7B scale.
  • The sources describe chat-only access, so self-hosting needs a confirmed weights release. TII presents the launch as "sovereign AI" capability (Arabian Reseller).
Read the full section
  • Dialect is a separate requirement. For UAE consumer-facing chat, customer service or government services, strong Arabic benchmark scores do not guarantee replies in the dialect. Teams should measure the register of replies separately from accuracy.
  • A replicable pattern. TII's documented approach (native web text, culture-focused MSA, glossary-constrained synthetic data, a native-speaker benchmark) could be applied to other low-resource dialects at 7B scale.
  • Procurement. The sources describe chat-only access, so self-hosting needs a confirmed weights release. TII presents the launch as "sovereign AI" capability (Arabian Reseller).

Sources

Read the full section
FOLLOW THE EVIDENCE

The source trail.

Sources (9)
A LITTLE LESS NOISE. A LOT MORE CONTEXT.

Stay curious.
Follow the evidence.

Independent perspectives, the original sources, and room for the questions that don't have easy answers.

How we build the brief