UAE & GCC4 min read

TII launches Falcon-Emirati alongside Arabic speech and OCR models in Abu Dhabi

Abu Dhabi's TII shipped three Arabic-first models: a 7B dialect-tuned LLM, a 1.6B speech recogniser and a 270M document model, alongside a native Emirati benchmark.

What happened

Summary of reporting by Technology Innovation Institute

On 6 October at Ai Everything Abu Dhabi, the Technology Innovation Institute (TII, part of ATRC) released three models focused on Arabic as spoken and written in the UAE. Falcon-Emirati is a 7-billion-parameter language model tuned for Emirati Arabic; Falcon-ASR is a 1.6-billion-parameter speech-to-text model covering Emirati Arabic, Modern Standard Arabic, English, French, Spanish and Portuguese; and Falcon-OCR-Arabic extracts text and structure from Arabic documents and images. All three are available via TII's Falcon Chat platform.

Falcon-Emirati builds on the Falcon-H1-Arabic family, a hybrid architecture combining Mamba state-space blocks and Transformer attention with context windows up to 256K tokens. TII adapted it with native Emirati web text, Modern Standard Arabic records covering Emirati heritage, and synthetic text constrained by Emirati glossaries. On Alyah, an evaluation set of 1,173 hand-collected multiple-choice items created with native speakers, it scored 84.83%, surpassing Jais-2-8B-Chat (78.09%), ALLaM-7B-Instruct-preview (77.24%), gemma-3-27b-it (74.68%) and Fanar-2-27B-Instruct (66.15%). In open-ended register evaluation, the model replied in Emirati dialect 52.1% of the time, compared to 5.3% for the next best competitor.

The accompanying speech and document models target compact efficiency. Falcon-ASR averaged a 20.92 word error rate across Arabic evaluation sets and outperformed a 30-billion-parameter multimodal system on TII's internal Emirati audio benchmark (22.73 WER against 26.80), delivering word-level timestamps. Falcon-OCR-Arabic achieved 81.9% text accuracy on 11,974 real-world samples, placing second to Gemini 3.5 Flash across a 17-model comparison and ranking first on official records, administrative forms, receipts and invoices. Reporting: TII announcement (https://www.tii.ae/insights/tii-launches-falcon-emirati-alongside-new-arabic-ai-models-speech-and-visual-text).

Read the original at Technology Innovation Institute

The Azrty take

Falcon-Emirati proves that dialect fidelity, not parameter volume, is the true quality gate for Arabic AI in the Gulf.

The decisive metric here is not 84.83% on Alyah. It is 52.1% against 5.3%. TII's LLM-judge study evaluated five models on 1,173 Emirati questions open-endedly to measure dialect fidelity: how often the response stayed in Emirati rather than slipping into Modern Standard Arabic or a fusha-English blend. Only Falcon-Emirati maintained dialect reliably: 52.1% against 5.3% for ALLaM-7B-Instruct-preview, 3.2% for gemma-3-27b-it, 2.0% for Jais-2-8B-Chat and 0.4% for Fanar-2-27B-Instruct. For GCC organisations running Arabic contact centres, government portals or field-service chat, this is the core issue. A citizen or customer in Abu Dhabi who hears MSA hears bureaucratic prose, not a human agent. Arabic pilots evaluated solely on MSA test sets will pass functional checks yet fail end users because the model responds accurately in the wrong register.

Model scale does not guarantee dialect capability. On the multiple-choice benchmark, Fanar-2-27B-Instruct (27B) scored 66.15% and gemma-3-27b-it (27B) scored 74.68%, falling behind Falcon-Emirati's 84.83% at 7B. Fanar-2-27B declined to answer on 26.2% of open-ended queries while the other systems stayed under 5%. The public Alyah benchmark confirms this across 53 models: the highest non-Falcon result was falcon-h1-arabic-7b-instruct at 82.18%, while Llama-3.3-70B-Instruct reached 69.74% and Qwen2.5-72B-Instruct achieved 74.6%. This reflects training data curation over raw compute, detailed in TII's Falcon-Emirati technical blog: native Emirati web content, MSA historical corpora, and synthetic generation anchored by Emirati glossaries so local dialect is not diluted. The foundational model is equally critical: Falcon-H1-Arabic, released in 3B, 7B and 34B parameter variants, combines Mamba state-space blocks with Transformer attention, scoring 75.36% on the Open Arabic LLM Leaderboard at 34B to outpace Qwen2.5-72B and Llama-3.3-70B on regional text.

The two companion releases have immediate operational utility. Falcon-ASR (1.6B) averaged 20.92 WER and 8.79 CER across evaluation sets on the Open Universal Arabic ASR Leaderboard, beating Audar-ASR-V1-turbo (23.17 at 2.35B), Cohere Transcribe Arabic (25.87) and omniASR LLM (28.32 at 7B). On TII's internal Emirati audio benchmark, it recorded 22.73 WER against 26.80 for Qwen3-Omni-30B-A3B-Instruct. Six supported languages including spoken Emirati, alongside word-level timestamps, provide an efficient contact-centre transcription layer. Falcon-OCR-Arabic (270M) is an equally significant efficiency win: 81.9% text accuracy across 11,974 document samples, ranking behind Gemini 3.5 Flash (84.3%) but ahead of Claude Opus 5.5 (79.2%) and GPT Astra (75.0%), while taking first place on official forms (84.2%), administrative templates (75.5%), receipts (71.6%) and invoices (72.5%). Built on the Falcon-OCR model card architecture (270M early-fusion Transformer), detailed in the Falcon Perception research paper and hosted in the Falcon-Perception repository with a vLLM container, it demonstrates that a specialised 270M model can outperform frontier models on regional administrative paperwork.

How Azrty would approach this technically: start with the 7B model rather than 34B. The hybrid Mamba and attention architecture with 256K context limits memory overhead, allowing organisations to handle extended dialogues on a single GPU node. Deploy the model behind a unified gateway to manage routing, budgets and access controls centrally (the architecture implemented by our FastLLM Proxy across local servers and hosted providers), while isolating the OCR service in its own container (ghcr.io/tiiuae/falcon-ocr:latest) so document parsing never contends with interactive chat for GPU memory. Before deploying to production, validate candidate checkpoints against the public Alyah harness using TII's LightEval fork:

bash\n# Acceptance gate: score your Arabic candidate on Alyah before any pilot\ngit clone https://github.com/amztheorytii/lighteval_em.git && cd lighteval_em\npip install -e .[multilingual] && pip install language_data\naccelerate launch --multi_gpu --num_processes=8 -m lighteval accelerate \\\n model_name=local-candidate,batch_size=8 'alyah' \\\n --output-dir results --load-tasks-multilingual --save-details\n

Incorporate a dialect-fidelity judge into this evaluation suite: configure an independent LLM to score responses for Emirati dialect register alongside semantic correctness, flagging regressions when answers drift into formal MSA. Common mistakes include procuring Arabic capabilities from global providers without local validation, evaluating exclusively on MSA text, and allocating large GPU clusters for 27B models when a 7B model handles regional dialect more effectively. The strategic advantage in the Gulf lies in an integrated local stack: regional dialect chat, speech transcription, and Arabic document processing that operate entirely within local sovereign infrastructure. Note the deployment boundary: Falcon-Emirati and Falcon-ASR currently run through Falcon Chat and hosted Spaces, whereas Falcon-OCR provides public weights. Teams planning private infrastructure deployments should prototype against the hosted APIs while preparing deployment pipelines for open weights releases.

What to do now

  1. Establish an Emirati dialect baseline: clone the Alyah harness (https://github.com/amztheorytii/lighteval_em) and evaluate your current Arabic production model against the 1,173-question benchmark. Treat scores below 77% as indicative of missing dialect comprehension rather than minor variance.
  2. Benchmark Falcon-OCR-Arabic on a sample of 200 local invoices, forms and official records using the Falcon-Perception vLLM container (ghcr.io/tiiuae/falcon-ocr:latest), measuring character accuracy against existing extraction tools using TII's 81.9% overall accuracy as a target.
  3. Run 20 to 30 recorded Emirati Arabic customer service calls through Falcon-ASR on Falcon Chat, calculating WER against your deployed transcription provider to see if you can approach TII's 22.73 WER reference benchmark.
  4. Implement a dialect-fidelity automated test in your continuous integration pipeline: deploy an LLM judge that evaluates dialect register independently from factual correctness on every prompt or model update.
Run local Arabic AI on your own infrastructureAzrty designs and operates dedicated platforms for regional LLMs, model gateways and vision pipelines across the GCC.
FalconArabic AILLMSpeech-to-TextAbu DhabiOpen Source

More from the Brief

TII launches Falcon-Emirati alongside Arabic speech and OCR models in Abu Dhabi: the Azrty take | Azrty Brief