HuggingFace Unveils Arabic LLM Leaderboard
Key point
The Open Arabic LLM Leaderboard has been released to professionally evaluate the performance of Arabic LLMs.
Details
The Open Arabic LLM Leaderboard (OALL) is a specialized benchmark platform designed to promote research and development in Arabic natural language processing (NLP).
Key benchmark datasets include the following:
- AlGhafa: Includes 22 datasets evaluating reading comprehension, sentiment analysis, question answering, and more.
- ACVA and AceGPT: Utilizes 58 datasets and Arabic-translated versions of MMLU and EXAMS to reflect linguistic complexity.
Technical configuration uses the lighteval library to perform evaluations, with the backend running on the TII cluster. Normalized log likelihood accuracy is used as the evaluation metric to fairly measure model performance.
Future plans include a leaderboard for RAG (retrieval-augmented generation) scenarios, a chatbot arena based on ELO scores, and the development of the OpenDolphin benchmark, which will include around 50 datasets.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.