Observed arrival · 2026-10-09
Mizan puts Arabic-language models on the scales
An open benchmark compares large language models’ Arabic abilities through blind reader judgments and a separate proficiency test marked as coming soon.
- For
- Researchers evaluating Arabic-language large language models
- Worth noticing
- The blind comparison uses Bradley–Terry estimates with 95% confidence intervals; the 150-question proficiency test is marked as coming soon.
Field notes
The comparison asks readers to choose between two answers without revealing which models produced them, or to mark the pair equivalent. The homepage reports 26 comparison questions and 152 human judgments, and explains that overlapping 95% confidence intervals result in tied ranks. A separate 150-question test is described but labeled as coming soon; its planned coverage includes dialects, grammar, spelling and diacritics, and open-ended writing.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue