AI Briefing
KO

Cleaning Up Arabic Benchmarks

·2026.04.21 19:09

Key point

QIMMA first validates Arabic benchmarks before evaluating LLMs on them.

1 / 2

Details

QIMMA(قِمّة) is a leaderboard that re-scores Arabic LLMs by validating benchmark quality before evaluation.

  • It unifies 109 subsets from 14 source benchmarks, totaling over 52K samples.
  • It covers 7 domains: culture, STEM, law, medicine, safety, literature, and coding.
  • 99% of the content is native Arabic, and it's the first Arabic leaderboard to include code evaluation.

Quality validation proceeds through a 2-stage pipeline.

  • Stage 1: Qwen3-235B-A22B-Instruct and DeepSeek-V3-671B each score samples on a 10-point rubric.
  • If either model scores below 7/10, the sample moves to Stage 2.
  • Stage 2: Native Arabic speakers conduct a final review of cultural context, dialectal nuances, subjective interpretation, and subtle errors.

The validation revealed recurring quality issues in existing benchmarks.

  • ArabicMMLU: 436 discarded out of 14,163, 3.1%
  • MizanQA: 41 discarded out of 1,769, 2.3%
  • PalmX: 25 discarded out of 3,001, 0.8%
  • MedAraBench: 33 discarded out of 4,960, 0.7%
  • FannOrFlop: 43 discarded out of 6,984, 0.6%

The issue types are categorized as answer quality errors, text/formatting corruption, cultural bias, and mismatches between gold answers and protocol.

For the code benchmarks 3LM HumanEval+ and 3LM MBPP+, no samples were discarded; only the Arabic problem descriptions were refined.

  • 88% of HumanEval+ prompts revised
  • 81% of MBPP+ prompts revised

Evaluation was made reproducible using LightEval, EvalPlus, and FannOrFlop, with different metrics applied for MCQ, QA, and code types. Results are public as of April 2026, and among the top models, Qwen/Qwen3.5-397B-A17B-FP8 ranked first with an average of 68.06.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.