//abdu.codes

Does Modern Standard Arabic Dominate Arabic Dialects in LLMs?

EMNLP 2026, Main Conference, short paper.

Separability by layer, mGPT

00.5116121824 NMI Layer Distinct languages Arabic varieties
The headline result. Distinct languages dip and recover. Arabic varieties flatten and stay there. No pivot, no dominance.

Large Language Models (LLMs) answer dialectal Arabic in Modern Standard Arabic (MSA). The usual explanation is that MSA dominates their internal space. Across 26 Arabic varieties and five models, it does not.

26Arabic varieties: 25 city dialects and MSA
5models, up to 8B parameters
0.1–0.2mean NMI, far below distinct languages
1/26uniform baseline, where dominance stays

The question

LLMs tend to answer in MSA even when prompted in a dialect. The explanation usually offered is representational: MSA is said to dominate the model's internal space, acting as a hidden pivot that dialect inputs are routed through.

That is a claim about geometry, not about output, and it had not been tested directly. If it were true, dialect representations should align systematically with MSA. We checked whether they do.

Method

We adapted the GMM-based language-dominance framework of Shani and Basirat (2025) from distinct languages to a cross-dialectal setting. The four steps, in order:

  • Separability is measured with Normalized Mutual Information (NMI) between variety labels and clusters at each layer. Low NMI means the varieties overlap.
  • Dominance is measured with a token-level likelihood ratio Λ comparing how strongly a token aligns with its own variety against each competing one. Dominance would appear as a directional pull toward one variety.

The test is symmetric: it is computed for every dialect pair at every layer, without assuming in advance which variety would dominate. Five models were evaluated on TU Dresden's Capella HPC cluster: BLOOM, mGPT, Qwen2.5-7B, Llama-3.1-8B and ArabianGPT-0.8B.

Findings

1

No variety dominates, MSA included.

Dominance scores sit at the uniform baseline of 1/26 across layers, and the share of tokens aligning better with a competing variety stays near 30% at every depth, rather than spiking in the middle layers as it does for distinct languages.

Figure 1Separability by layer. Arabic varieties (red) stay low and flat where distinct languages (navy) recover, and the low range extends one to three layers further. Switch model, hover for values, click a key to mute a series.
Figure 2Tokens that fit a competing variety better than their own stay near 30% at every layer. The absence of a middle-layer spike indicates ambiguity rather than a directional pull. Hover for per-layer values.
2

Dialect identity survives the overlap.

Low separability is not collapse. The confusion matrices keep clear diagonals and a supervised linear probe recovers dialect labels well above chance. The varieties are densely packed, not erased.

3

Where compression happens depends on the architecture.

Separability is stable in Llama-3.1, drops early in Qwen2.5 and collapses late in Arabic-native ArabianGPT. Training predominantly on Arabic does not make dialects more separable; it shifts where the compression occurs.

Figure 3Each architecture on its own scale, as in the paper: stable (Llama-3.1), an early drop (Qwen2.5) and a late collapse (ArabianGPT). Llama's whole curve spans 0.018 NMI, so a shared axis would flatten it. Hover any panel for per-layer values.

Why it matters

MSA-biased generation is real, but it is not explained by MSA dominating the hidden states. Output preference and representational geometry are separate phenomena, and analyses of multilingual and dialectal LLMs should treat them separately. If the bias is not internal routing, it plausibly originates elsewhere: the share of MSA in pretraining and instruction-tuning data, its more standardised orthography, or decoding-time preferences.

Limitations

MADAR covers the travel domain only, with 400 sentences per dialect for BLOOM and mGPT and 100 for the other models, so the results are exploratory rather than definitive. The ablations use rule-based POS tagging. Robustness checks with POS filtering, sentence-prefix removal and varied PCA dimensionality preserve the trends.

Citation

BibTeX
@inproceedings{sallouh2026msa,
  title     = {Does Modern Standard Arabic (MSA) Dominate Arabic Dialects in LLMs?
               A Representation-Level Analysis},
  author    = {Sallouh, Abdu and Popovi{\v{c}}, Nicholas and F{\"a}rber, Michael},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods in
               Natural Language Processing (EMNLP)},
  year      = {2026}
}