Best Arabic LLM for business: Jais 2, Falcon or GPT-5?
For a retailer answering customers in Gulf Arabic, the model choice matters more than the chatbot vendor. Two UAE-built models, Jais 2 and Falcon-H1-Arabic, are now production-grade alternatives to GPT-5 and Claude. Here is how they compare on dialect, cost and where the data is processed, with the questions to ask a vendor and a test you can run this month.
The best Arabic LLM for business is no longer a global API by default. For dialect-heavy customer support, high message volumes, or regulated UAE workloads, UAE-built open models Jais 2 and Falcon-H1-Arabic are the strongest options, while frontier APIs such as GPT-5 or Claude suit low-volume MSA and mixed English traffic without strict residency rules. If your shop answers customers in Gulf Arabic, the single biggest quality decision is not which chatbot vendor you sign with. It is which language model sits underneath the chatbot. In the last year that choice changed materially: on 9 December 2025 G42's Inception, MBZUAI and Cerebras released Jais 2, an 8B and 70B Arabic-centric model trained from scratch under Apache 2.0 (PR Newswire), and on 5 January 2026 Abu Dhabi's Technology Innovation Institute released Falcon-H1-Arabic in 3B, 7B and 34B sizes (TII announcement). Both are open-weight, both run on your own infrastructure in the UAE, and both beat general-purpose global models on Arabic benchmarks. This article compares them with the big API models on dialect, cost and data residency, gives you ten questions to ask a chatbot vendor, and walks through a test you can run this month on your own customer messages.
What is the best Arabic LLM for business?
The best Arabic LLM for business in 2026 is not one model. It is a routing decision:
- Dialect-heavy, high-volume or regulated Arabic traffic (retail customer service in the UAE, banking, government services): run Jais 2 or Falcon-H1-Arabic on your own GPU servers, inside the UAE. Start with Jais-2-70B if you have two H100-class GPUs, Falcon-H1-Arabic-7B if you have one.
- Low-volume, mixed English traffic with no residency constraint: a global API model behind an OpenAI-compatible gateway. GPT-5 costs $1.25 per million input tokens and $10 per million output tokens (OpenAI pricing), so at small scale the API bill is trivial and quality on Modern Standard Arabic (MSA) is strong.
- Most UAE retailers we work with end up hybrid: Arabic requests route to a self-hosted Arabic model, English requests route to a global API, all through one gateway so you can swap either side later.
And one rule that overrides everything above: the model that scores best on your 100 real customer messages wins, regardless of what any leaderboard says. The test is cheap and this article includes it.
Why does the model matter more than the chatbot vendor?
Every Arabic customer-service stack has the same four layers: a channel (WhatsApp Business, web chat, voice), orchestration (intents, tools, escalation to a human agent), retrieval (your order system, returns policy, delivery zones) and the model that actually writes the Arabic reply. The first three layers are commodities now. Fine-tuning an intent router is a solved problem, and any competent team can wire retrieval into a retailer's order API.
The model layer is where Arabic deployments come apart. Arabic is diglossic: customers write in Gulf dialect, Arabizi (Arabic in Latin script, like "mt2kher 3n el order mta3i"), code-switched Arabic-English, and occasionally MSA, while most documentation and policies are written in MSA. A model trained mostly on English and formal Arabic handles the MSA fine and stumbles on the dialect, which shows up in your metrics as customers abandoning chats, repeating themselves, or asking for a human.
Until recently the workaround was a translation layer: translate Arabic to English, answer with an English model, translate back. That cost latency and quality, and it sent the raw conversation text abroad. The 2026 releases make it obsolete, and consultants running sovereign AI engagements in the GCC now treat Jais 2 and Falcon-H1-Arabic as the default for Arabic workloads rather than research curiosities (Codenovai comparison).
Which UAE-built Arabic models are available?
Jais 2: the biggest open Arabic model, trained from scratch
Jais 2 comes from Inception (a G42 company), the Institute of Foundation Models at MBZUAI and Cerebras, released in 8B and 70B sizes, both trained from scratch on Arabic and English rather than adapted from an English base (Jais 2 technical report). The facts that matter for a business buyer:
- Licence: Apache 2.0, confirmed on the Hugging Face model card. The cleanest possible answer for a legal team reviewing an LLM procurement.
- Dialect coverage: designed for MSA, regional dialects and mixed Arabic-English code-switching, including Arabizi. The paper reports broad dialectal and script coverage in the training corpus.
- Benchmarks: on the Open Arabic LLM Leaderboard 2 (OALL2) suite, Jais-2-8B scores a 72.40% macro-average, the best of all models at 13B parameters and below, ahead of Fanar-1-9B-Instruct (68.97%), ALLaM-7B-Instruct-preview (67.29%) and Llama-3.1-8B-Instruct (51.60%). Jais-2-70B scores 79.36%, ahead of Llama-3.3-70B-Instruct (74.23%) and Qwen2.5-72B-Instruct (71.20%). It also leads on AraGen generative tasks and on culturally grounded domains, and scores well on the AraTrust safety benchmark.
- Vocabulary: a custom Arabic-centric vocabulary of 150,272 tokens. This is not cosmetic: English-centred tokenisers split Arabic words into more tokens, which inflates both latency and cost per conversation on global APIs.
- The catch: an 8,192-token context window. Fine for a customer-service turn with a retrieval chunk. Tight for long RAG contexts, long policy documents or multi-turn conversations that carry history.
The 70B model is, to the authors' knowledge, the largest open Arabic-centric LLM trained from scratch, and it is deployed as a public chat app on Cerebras hardware serving up to 2,000 tokens per second (Cerebras blog). If a national-scale model runs it in production chat, it will survive your retail peak season.
Falcon-H1-Arabic: smaller, faster, with a 256K context window
TII's Falcon-H1-Arabic, announced 5 January 2026, takes a different route (TII technical blog):
- Architecture: a hybrid Mamba-Transformer (Falcon-H1) design where state-space and attention layers run in parallel in every block. Practically, it means linear-time behaviour on long sequences, which is how the family gets 128K context at 3B and 256K context at 7B and 34B. For retrieval-heavy Arabic assistants that need to hold a large policy document in context, this is the differentiator.
- Sizes: 3B, 7B, 34B. The 7B is aimed squarely at production chat; the 34B is the flagship. A 34B model at BF16 needs roughly 68GB for weights, so it fits on a single 80GB H100; the 7B needs about 14GB and runs on a single modern 24GB card. Jais-2-70B needs around 140GB for weights, so two H100 80GB cards or one 141GB H200.
- Benchmarks: 71.7% on the Open Arabic LLM Leaderboard (OALL) for the 7B, beating Fanar-9B, ALLaM-7B and Qwen3-8B; roughly 75% for the 34B, ahead of Llama-3.3-70B and AceGPT2-32B. On 3LM, the Arabic STEM benchmark, the 34B reaches about 96% on the native split. TII publishes 7B AraDiCE dialect scores in the mid-50s and ArabCulture near 80%.
- Data: trained on around 300 billion tokens in a near-equal mix of Arabic, English and multilingual content, with expanded dialectal sources for Gulf, Levantine, Egyptian and North African Arabic, and no reliance on machine translation (TII product page).
- Licence: the Falcon family ships under the TII Falcon-LLM licence, which permits commercial use subject to attribution and acceptable-use conditions. Fine for almost every deployment, but less clean than Apache 2.0 if your lawyers are strict about open-source pedigree.
- Tooling: runs on Hugging Face transformers, vLLM (version 0.9.0 or later) and llama.cpp per the Falcon-H1 model cards.
For completeness: two other sovereign models matter in the GCC. Saudi Arabia's ALLaM (SDAIA, now run by HUMAIN, in 7B to 70B sizes, powering the HUMAIN Chat app on its 34B model) and Qatar's Fanar from QCRI (7B to 27B, with Fanar 2.0 multimodal). If your operation is Saudi-headquartered or Qatar-headquartered, both belong on your shortlist; for UAE deployments the local pair is usually the default (Arabic LLM landscape overview).
How do Arabic models compare with GPT-5 and Claude?
The benchmark picture, with one honest caveat: Jais 2's scores come from the OALL2 suite in the technical report; TII reports Falcon-H1-Arabic on OALL with vLLM as the backend. The suites overlap heavily (both descend from the Open Arabic LLM Leaderboard, which notably rebuilt itself around natively written Arabic questions after dumping translated-from-English tests) but they are not identical, so treat this table as directional and run your own eval before you commit.
| Model | Params | OALL macro-average (as reported) | Context | Weights |
|---|---|---|---|---|
| Jais-2-70B | 70B | 79.36% (OALL2) | 8,192 | Open, Apache 2.0 |
| Jais-2-8B | 8B | 72.40% (OALL2) | 8,192 | Open, Apache 2.0 |
| Falcon-H1-Arabic-34B | 34B | ~75% (OALL) | 256K | Open, Falcon-LLM licence |
| Falcon-H1-Arabic-7B | 7B | 71.7% (OALL) | 256K | Open, Falcon-LLM licence |
| Llama-3.3-70B-Instruct | 70B | 74.23% (OALL2) | 128K | Open, Llama licence |
| Qwen2.5-72B-Instruct | 72B | 71.20% (OALL2) | 32K | Open |
| ALLaM-7B-Instruct | 7B | 67.29% (OALL2) | 4K | Partial |
| Fanar-1-9B-Instruct | 9B | 68.97% (OALL2) | 32K | Open |
The business-level comparison:
| Attribute | Jais 2 | Falcon-H1-Arabic | Global APIs (GPT-5, Claude) |
|---|---|---|---|
| Origin | UAE (Inception, MBZUAI, Cerebras) | UAE (TII) | US |
| Sizes | 8B, 70B | 3B, 7B, 34B | Frontier scale |
| Context window | 8,192 | 128K to 256K | 400K to 1M |
| Gulf dialect + Arabizi | Strong (built for it) | Strong (expanded dialect corpora) | Adequate on MSA, weaker on dialect |
| Arabic benchmarks | Best open scores at both sizes | State of the art per parameter class | Not top of Arabic boards |
| Where inference runs | Your infrastructure | Your infrastructure | Provider regions, mostly outside the UAE |
| Cost model | GPU capex or rental, predictable | GPU capex or rental, cheapest at 3B/7B | Per token, scales with volume |
| Fine-tuning | You own the weights | You own the weights | Limited or impossible |
| English, code, hard reasoning | Competitive, not frontier | Competitive, not frontier | Best in class |
| Deprecation risk | You control the weights | You control the weights | Provider retires models |
What do global models still do better?
None of this means GPT-5 or Claude are the wrong choice. Verified list prices per million tokens: GPT-5 at $1.25 input and $10.00 output (OpenAI); Claude Sonnet 4.6 at $3 input and $15 output with a 1M-token context window, Claude Sonnet 5 at $2 and $10, Claude Haiku 4.5 at $1 and $5 (Anthropic pricing).
The honest split, consistent with what practitioners report from GCC deployments (Codenovai):
| Task | Arabic models (Jais 2, Falcon) | Global frontier models |
|---|---|---|
| MSA Q&A over retrieved context | Strong | Strong |
| Gulf dialect customer support | Strongest | Adequate |
| Arabic-English code-switching | Strong | Adequate |
| Arabic cultural and religious context | Strong | Adequate |
| English-only tasks, code, hard maths | Competitive | Strongest |
| Structured JSON output, tool calling | Strong | Strongest |
In other words: for the narrow job of answering a customer in Gulf Arabic, the Arabic models match or beat the frontier. For everything around it (agentic workflows, tool use, English content), the frontier still leads, which is precisely why hybrid routing exists.
How much does running an Arabic LLM cost?
Take a mid-sized UAE e-commerce retailer, call it GulfMart, with a WhatsApp and web-chat assistant handling Arabic support.
Assumptions: 12,000 Arabic conversations per month, four turns each, roughly 400 input tokens per turn (system prompt, a retrieval chunk, the message) and 200 output tokens. That is 1,600 in and 800 out per conversation, or 19.2M input and 9.6M output tokens per month.
Option 1, GPT-5 API: 19.2M × $1.25 = $24 input, 9.6M × $10 = $96 output. Total $120 per month, about AED 440 at the pegged rate of 3.67.
Option 2, Claude Sonnet 4.6 API: 19.2M × $3 = $57.60 input, 9.6M × $15 = $144 output. Total $202 per month, about AED 740.
Option 3, self-hosted Falcon-H1-Arabic-7B: one modern GPU. Even before counting GPU rental, at GulfMart's volume the API option wins on pure cost. Two caveats keep the maths honest: real deployments burn two to ten times the naive token count through agent loops, retries, long conversation history and Arabic's token overhead on English-centred tokenisers; and a dedicated GPU carries its own monthly cost that does not shrink when traffic does.
So when does self-hosting pay? One consultancy publishing its GCC deployment economics puts a cloud-only Claude deployment at AED 35,000 to 80,000 per month at roughly 3M tokens per day, a sovereign Falcon-H1 70B tier at AED 280,000 capex plus AED 12,000 to 18,000 monthly opex, and a break-even between 1.5 and 2M tokens per day (Codenovai). In our experience the break-even point moves with your token intensity, but the order of magnitude holds: past a few million tokens per day, owned Arabic inference is cheaper than frontier APIs, and below it the decision is not about cost at all.
What flips the decision before cost does is compliance and control. The same source notes the UAE Central Bank's sovereign financial cloud (launched February 2026) makes data residency a hard constraint for in-scope workloads, and Dubai's agentic AI requirements favour models you can fully audit, with traceable provenance. Add three business reasons that apply to any retailer: customer transcripts never leave your infrastructure; you can fine-tune on your own season of conversations and keep the weights; and no vendor can deprecate your model from under you.
What is the best architecture for Arabic customer support?
The pattern that works in practice is one OpenAI-compatible gateway with two model pools and a language check at the entry point:
# schematic gateway routing: Arabic stays on UAE infrastructure
routes:
- match: lang == "ar" # detected at request entry
pool: arabic # Falcon-H1-Arabic or Jais 2 on your GPU servers
fallback: arabic_small # Falcon-H1-Arabic-3B for spikes
- match: lang == "en"
pool: global # GPT-5 or Claude, if data classification allows
fallback: arabic
- match: lang == "mixed" # code-switched, Arabizi
pool: arabic
budgets:
per_customer_monthly_usd: 50
The gateway matters more than it looks: it makes every model an OpenAI-compatible endpoint, so swapping Falcon for Jais, or adding a global provider for a new English market, is a config change and not a rewrite. We run client stacks this way with FastLLM Proxy, which puts routing, budgets and access control for your own LLM servers and hosted providers in one place, and it is how we build AI infrastructure for teams that want the Arabic layer on their own GPUs.
Questions to ask a vendor, and the test to run before you sign
Ten questions for a chatbot vendor
- Which model, exactly, answers my Arabic customers? A named model and version, not "AI-powered".
- Does it handle Gulf dialect and Arabizi, and can you show me the evaluation behind that claim?
- Where does inference run: which region, whose tenant? Where do transcripts sit, for how long, and are they used for training?
- Can we swap the model later through an OpenAI-compatible endpoint, without re-platforming?
- Who owns the conversation data and any fine-tuned weights if we leave?
- How are Arabic tokens counted? (English-centred tokenisers inflate Arabic cost.)
- What are the safety guardrails in Arabic, not just English, and how are refusals handled in dialect?
- What happens to our service and price when the underlying model is deprecated or repriced?
- Will you run an evaluation on 100 of our own messages before we sign?
- What is your p95 latency per turn, and your measured cost per resolved conversation?
If a vendor cannot answer 3, 4 and 9, keep shopping.
The test: run it yourself, this month
Benchmarks are directional; your customers decide. The procedure:
- Pull 80 to 100 real customer messages from the last month: WhatsApp, live chat, email. Anonymise names and order numbers. Keep the dialect mix as it actually is, including Arabizi and code-switching.
- Write the expected answer for each, or a short rubric (correct policy applied, dialect-appropriate tone, no invented order numbers, human escalation offered when appropriate).
- Run every candidate model on the identical prompt and identical retrieved context. Use the same OpenAI-compatible harness for all of them, so the only variable is the model:
from openai import OpenAI
MODELS = {
# any OpenAI-compatible endpoint: provider API, gateway, or your own vLLM server
"gpt5": dict(base_url="https://api.openai.com/v1", model="gpt-5"),
"sonnet": dict(base_url="https://your-gateway/v1", model="claude-sonnet-4-6"),
"falcon": dict(base_url="http://localhost:8000/v1", model="falcon-h1-arabic-7b"),
"jais": dict(base_url="https://your-gateway/v1", model="jais-2-8b"),
}
CASES = [\n # {"id": "001", "msg": "<real anonymised customer message>", "rubric": "..."},\n]
for name, cfg in MODELS.items():
client = OpenAI(base_url=cfg["base_url"], api_key="REDACTED")
for case in CASES:
resp = client.chat.completions.create(
model=cfg["model"],
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": case["msg"]},
],
)
print(name, case["id"], resp.usage.prompt_tokens, resp.usage.completion_tokens)
# save outputs for blind scoring; log token counts for the cost column
- Have two native Gulf Arabic speakers score the outputs blind, without knowing which model produced which, and tally by category: policy correctness, dialect quality, Arabizi handling, safety refusals.
- Include at least three cases that stress dialect specifically, for example:
- MSA: "أين طلبي رقم 84512؟ مر عليه خمسة أيام." (Where is order 84512? Five days have passed.)
- Gulf dialect: "شحال يستغرق التوصيل لأبوظبي؟ والطلب عندي متأخر من أمس." (How long does delivery to Abu Dhabi take? My order is late already.)
- Arabizi: "salam, mt2akher 3an el order mta3y, momken tshofooly wen wsel?" (Hi, my order is late, can you check where it is?)
- Decide with a simple rule: the model with the highest policy correctness at acceptable dialect tone wins. Check token counts to compute the true cost per conversation. If the Arabic models win on quality but lose on convenience, that is the case for the self-hosted route; if a global API wins at your volume, start there and keep the gateway so you can switch when volume or regulation forces the issue.
Budget one to two weeks for a single use case. That is the difference between a defensible model selection and a guess from a leaderboard.
What to do next
- This week: pull 100 real customer messages and run the harness above against GPT-5, Claude, Falcon-H1-Arabic and Jais 2. Two weeks of effort, and it converts the biggest open question in your Arabic channel into evidence.
- Decide the residency question first, not last: if your transcripts are regulated or strategic, the answer is self-hosted Arabic models on UAE infrastructure regardless of token maths, and the only remaining question is Jais 2 versus Falcon-H1-Arabic on your corpus.
- Whatever you choose, put an OpenAI-compatible gateway in front of it now. Model prices and rankings move every quarter; a gateway turns each move into a config change instead of a project.
- When you are ready to size and build the self-hosted layer, that is the work we do in AI infrastructure: GPUs and models on your own infrastructure, one gateway for every model, run for you.
The models finally exist. The vendors will not test them for you. Run the test, then buy.
Link to this article
Citing this in your own writing? Use the permanent link below.https://www.azrty.com/blog/best-arabic-llm-for-business-jais-2-falcon-or-gpt-5
<a href="https://www.azrty.com/blog/best-arabic-llm-for-business-jais-2-falcon-or-gpt-5">Best Arabic LLM for business: Jais 2, Falcon or GPT-5?</a> (Azrty)