LMArena / Arena · Blind human preference
Arena
Pairwise, anonymous human votes compare model responses and media outputs generated from the same prompt.
What it measures
- •Which response a voter prefers in a blind head-to-head
- •Category and modality leaderboards derived from many pairwise votes
- •Confidence ranges where the leaderboard exposes them
Use it for
Finding models people tend to prefer for open-ended chat, style, instruction following, and media quality.
Continuously updated from new votes; use the live leaderboard and its current category filters.
Read with care
Preference is audience- and prompt-distribution dependent. Elo-like ratings are relative, close ranks can overlap, and popularity is not correctness or safety.
From evidence to usage
Run the closest available path
These are availability mappings, not claims that an app.nz route reproduces the source’s score or harness.
Recreate a blind-text shortlist
app/auto · frontier panelStart with a shared prompt, compare exact outputs, then repeat with your own acceptance criteria.
Compare image families
GPT Image · FLUX · SDXLRun proprietary hosted models beside open workflows rather than treating one global preference score as universal.
Make the final decision on your data
Keep public evidence as context, then score representative inputs with a frozen rubric. Inspect failures—not only the mean.