AI Pulse by Inblix

54 Arabic AI models fail basic Emirati dialect test, new 1,173-question benchmark reveals

Hugging Face Blog · Jan 27, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: 54 Arabic AI models fail basic Emirati dialect test, new 1,173-question benchmark reveals

If you’ve ever tried speaking to an AI in a language that isn’t textbook-perfect, you know the pain. A new benchmark called Alyah makes that failure official for Arabic, and the results are a wake-up call. The team behind it put 54 large language models to the test—not on formal Modern Standard Arabic, but on the Emirati dialect. The punchline? Even the best models are effectively lost in the souk.

Alyah, meaning “North Star” in Emirati, is a manually curated dataset of 1,173 multiple-choice questions. Forget simple vocabulary drills; this benchmark probes the soul of the dialect. Models are quizzed on culturally embedded greetings, oral poetry fragments, proverbs, and expressions where literal translation is useless. The source material came directly from native speakers, ensuring the kind of authenticity you can’t scrape from the web. Each question has four possible answers, with the wrong ones generated synthetically and then reviewed to make them genuinely tricky. The benchmark covers a spectrum of difficulty, from everyday phrases to dense heritage questions, and difficulty isn’t assigned by a person—it’s calculated coldly by how often the models actually crash and burn.

The evaluation lineup was massive: 23 base models and 31 instruction-tuned variants, including regional heavyweights like Jais and Allam, plus global players like LLaMA and Qwen. The scoring prioritized semantic understanding, not just keyword matching, which is crucial when a single greeting can have a dozen valid, culturally correct responses. The study’s design exposes a brutal truth: instruction tuning, often hailed as a fix for making models more helpful, doesn’t magically imbue them with deep cultural intuition. The categories where models consistently faceplanted aren’t just harder—they reveal a systemic blindness to pragmatic meaning that formal language benchmarks completely miss.

This isn’t just about a single dialect. It exposes a foundational crack in how we evaluate AI for most of the world. A model that aces a newswire test can still sound like a clueless tourist when a real conversation starts. The Alyah benchmark draws a clear line in the sand: reading a culture’s formal documents is not the same as understanding its people. For any developer hoping to deploy an AI assistant in the Gulf, these numbers are a direct indictment of the training data status quo.

💡 Key Takeaways

  1. Even Arabic-native LLMs like Jais and Allam struggle significantly with the Emirati dialect, proving that regional language support doesn't equal dialect fluency.
  2. The benchmark's difficulty isn't subjective; it's defined by consistent model failure, highlighting a systematic inability to interpret culturally embedded expressions and poetry.
  3. Instruction-tuned models showed no inherent advantage over base models in cultural understanding, suggesting alignment techniques don't fix gaps in pragmatic, non-formal training data.
  4. Manual data curation from native speakers was essential because the authentic expressions tested in Alyah are rarely documented online, meaning web-scraped training sets create a cultural blind spot.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles