All eval sources

Hugging Face community · Reproducible open-model evaluation

Hugging Face OpenEvals

A route into task-specific, open leaderboards and evaluation artifacts after the original Open LLM Leaderboard was retired.

What it measures

  • Task-specific open leaderboards rather than one permanent universal ranking
  • Public model, dataset, harness, and result artifacts where a leaderboard provides them
  • Open-weight models that can be downloaded and hosted independently

Use it for

Finding reproducible evidence for an open model and following it through to the exact model repository or weights.

The original Open LLM Leaderboard stopped accepting evaluations in 2025; use the maintained leaderboard directory.

Read with care

“Open” leaderboards vary in maintenance, task quality, licenses, quantization, prompts, and hardware. Inspect each leaderboard card and result artifact.

From evidence to usage

Run the closest available path

These are availability mappings, not claims that an app.nz route reproduces the source’s score or harness.

Make the final decision on your data

Keep public evidence as context, then score representative inputs with a frozen rubric. Inspect failures—not only the mean.

Build an eval set