07/20/2026
"There is no universally best ASR model for every Arabic dialect." That's Sanjika Hewavitharana, Head of AI Research at aixplain, on what this quarter's Dialectal Arabic ASR Benchmark makes clear.
For the July 2026 edition, we sat down with Sanjika Hewavitharana, Head of AI Research at aixplain, and Shreyas Sharma, ML Research Engineer at aixplain, the scientists behind the report, to talk through what changed. Two additions define this edition, and we moved fast on both.
Cohere Transcribe joins the lineup, bringing us to six systems. As soon as it was available, we put it head-to-head against the field. The result? It ranked second overall, behind only Azure, with best-in-class performance on Chami and Najdi.
UAE Arabic joins the dialect set, bringing us to five dialects and extending our coverage of Gulf Arabic.
A few things the data made clear:
"One model rarely wins everywhere. Cohere Transcribe excelled on Chami and Najdi, then dropped sharply on Hijazi." Said Sanjika, "A model that performs well on one regional dialect may struggle with another due to differences in pronunciation, vocabulary, or training data."
"Gulf Arabic" isn't one thing. Adding the UAE dialect revealed that even geographically close dialects behave differently. Najdi and UAE share much linguistically, yet models didn't perform identically on both. Treating Gulf Arabic as a single category reveals real performance gaps.
Cohere Transcribe is emerging as a compelling middle ground, with substantially better accuracy than Whisper across several dialects, without Azure's latency profile.
The takeaway holds across the board: evaluate models against your target dialects and operational requirements, not a single metric.
Full Report, link in bio.