AI Briefing
KO

STEM and Code Benchmark '3LM' Released for Arabic LLMs

·2025.08.01 23:25

Key point

A new benchmark called 3LM has been released to evaluate the STEM and code generation capabilities of Arabic LLMs.

1 / 2

Details

A new benchmark called 3LM(علم) has been announced to evaluate the technical capabilities of Arabic LLMs. Unlike existing benchmarks that have focused on general tasks such as summarization or sentiment analysis, 3LM focuses on verifying scientific reasoning and programming ability.

The benchmark consists of the following three datasets:

  • Native STEM: 865 real Arabic STEM multiple-choice questions (physics, chemistry, biology, mathematics, geography) extracted from 8th-12th grade textbooks and exam questions.
  • Synthetic STEM: 1,744 high-difficulty reasoning questions generated using the YourBench pipeline, testing conceptual and analytical thinking ability.
  • Arabic Code Benchmarks: A dataset translated and adapted into Arabic from HumanEval+ and MBPP+, evaluating code generation ability through Arabic prompts.

3LM secured data quality and logical accuracy through a process of OCR, LLM-based extraction, manual review, and backtranslation.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.