SyncAI.news, a Varaisys broadcasting
The Open Arabic LLM Leaderboard 2
HF

Hugging Face Blog

· 1 min read

AI LabsHugging Face Blog

The Open Arabic LLM Leaderboard 2

Current status of Arabic LLMs leaderboards

The growing availability of LLMs supporting Arabic, both as monolingual and multilingual models, prompted the community to create dedicated Arabic language leaderboards. Previously, Arabic-focused leaderboards were typically confined to narrow benchmarks introduced by specific authors, often as demos for their work. In these cases, the authors would set up leaderboards to demonstrate how models performed on a particular task or dataset. Alternatively, other leaderboards required users to run evaluations on their own computing resources and then submit a JSON file containing their results for display.

While these approaches helped spark initial interest in Arabic benchmarking, they also introduced several challenges:

  1. Resource Limitations: Many community members lack access to the substantial computational resources needed to evaluate all available open-source models in order to establish which model would be best for their downstream project or application, being forced to rely only on the results shared by model makers in their documentation, which many times does not allow for a direct comparison. This high cost in both time and compute power can become a major barrier to participation in further developing Arabic LLMs, making a leaderboard a valuable shared resource.
  2. Integrity of Reported Results: Because some platforms required users to evaluate their models independently and then simply submit a file of scores, there was no robust mechanism to ensure those results were accurate or even produced through a genuine evaluation. This lack of centralized verification could potentially undermine the credibility and fairness of the leaderboard.

Impact of the previous leaderboard

Why do we need a new leaderboard?

What's new in this version?

From the first version of the Open Arabic LLM Leaderboard (OALL), we keep the following benchmark datasets:

We enrich the leaderboard by adding the following datasets, released in the past year:

Original source

This story was published by Hugging Face Blog. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on huggingface.co

Similar News