Does Modern Standard Arabic Dominate Arabic Dialects in LLMs?
EMNLP 2026, Main Conference, short paper.
Large Language Models (LLMs) answer dialectal Arabic in Modern Standard Arabic (MSA). The usual explanation is that MSA dominates their internal space. Across 26 Arabic varieties and five models, it does not.
The question
LLMs tend to answer in MSA even when prompted in a dialect. The explanation usually offered is representational: MSA is said to dominate the model's internal space, acting as a hidden pivot that dialect inputs are routed through.
That is a claim about geometry, not about output, and it had not been tested directly. If it were true, dialect representations should align systematically with MSA. We checked whether they do.
Method
We adapted the GMM-based language-dominance framework of Shani and Basirat (2025) from distinct languages to a cross-dialectal setting. The four steps, in order:
- Separability is measured with Normalized Mutual Information (NMI) between variety labels and clusters at each layer. Low NMI means the varieties overlap.
- Dominance is measured with a token-level likelihood ratio Λ comparing how strongly a token aligns with its own variety against each competing one. Dominance would appear as a directional pull toward one variety.
The test is symmetric: it is computed for every dialect pair at every layer, without assuming in advance which variety would dominate. Five models were evaluated on TU Dresden's Capella HPC cluster: BLOOM, mGPT, Qwen2.5-7B, Llama-3.1-8B and ArabianGPT-0.8B.
Findings
No variety dominates, MSA included.
Dominance scores sit at the uniform baseline of 1/26 across layers, and the share of tokens aligning better with a competing variety stays near 30% at every depth, rather than spiking in the middle layers as it does for distinct languages.
Dialect identity survives the overlap.
Low separability is not collapse. The confusion matrices keep clear diagonals and a supervised linear probe recovers dialect labels well above chance. The varieties are densely packed, not erased.
Where compression happens depends on the architecture.
Separability is stable in Llama-3.1, drops early in Qwen2.5 and collapses late in Arabic-native ArabianGPT. Training predominantly on Arabic does not make dialects more separable; it shifts where the compression occurs.
Why it matters
MSA-biased generation is real, but it is not explained by MSA dominating the hidden states. Output preference and representational geometry are separate phenomena, and analyses of multilingual and dialectal LLMs should treat them separately. If the bias is not internal routing, it plausibly originates elsewhere: the share of MSA in pretraining and instruction-tuning data, its more standardised orthography, or decoding-time preferences.
Limitations
MADAR covers the travel domain only, with 400 sentences per dialect for BLOOM and mGPT and 100 for the other models, so the results are exploratory rather than definitive. The ablations use rule-based POS tagging. Robustness checks with POS filtering, sentence-prefix removal and varied PCA dimensionality preserve the trends.
Citation
@inproceedings{sallouh2026msa,
title = {Does Modern Standard Arabic (MSA) Dominate Arabic Dialects in LLMs?
A Representation-Level Analysis},
author = {Sallouh, Abdu and Popovi{\v{c}}, Nicholas and F{\"a}rber, Michael},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in
Natural Language Processing (EMNLP)},
year = {2026}
}