All eval sources

LMArena / Arena · Blind human preference

Arena

Pairwise, anonymous human votes compare model responses and media outputs generated from the same prompt.

What it measures

  • Which response a voter prefers in a blind head-to-head
  • Category and modality leaderboards derived from many pairwise votes
  • Confidence ranges where the leaderboard exposes them

Use it for

Finding models people tend to prefer for open-ended chat, style, instruction following, and media quality.

Continuously updated from new votes; use the live leaderboard and its current category filters.

Read with care

Preference is audience- and prompt-distribution dependent. Elo-like ratings are relative, close ranks can overlap, and popularity is not correctness or safety.

From evidence to usage

Run the closest available path

These are availability mappings, not claims that an app.nz route reproduces the source’s score or harness.

Make the final decision on your data

Keep public evidence as context, then score representative inputs with a frozen rubric. Inspect failures—not only the mean.

Build an eval set