A leaderboard is a lead,
not the decision.
Read independent evidence from Artificial Analysis and other maintained evals, open the exact models we host, then verify the shortlist on your prompts. One flow from public benchmark to production usage.
01
Find evidence
Use live sources, methodology, and uncertainty—not copied score snapshots.
02
Shortlist models
Match the benchmark modality and task distribution to the work you actually have.
03
Run the model
Open hosted gateway inference, a cloud agent, or a tested Comfy GPU workflow.
04
Keep the regression
Turn representative prompts and failures into a repeatable acceptance suite.
Independent evidence
Choose the question before the ranking
Every source below links to its live result and methodology. app.nz does not present old leaderboard values as current facts.
Quality, price, and measured speed
Artificial Analysis
Independent model analysis spanning language intelligence, price and latency plus blind preference arenas for image and video generation.
Blind human preference
Arena
Pairwise, anonymous human votes compare model responses and media outputs generated from the same prompt.
Fresh, objective language tasks
LiveBench
A frequently refreshed benchmark designed around recent questions, objective ground-truth scoring, and reduced test contamination.
Repository-level software repair
SWE-bench Verified
A human-validated 500-instance subset of real GitHub issues, evaluated by applying generated patches and running repository tests.
Reproducible open-model evaluation
Hugging Face OpenEvals
A route into task-specific, open leaderboards and evaluation artifacts after the original Open LLM Leaderboard was retired.
The useful eval loop
Public benchmarks are discovery data. Your acceptance suite is deployment data. Keep both, and record model version, provider route, prompt, settings, latency, cost, and failures.
Read the field guideStart wide
Use multiple source types: objective tasks, human preference, agentic work, speed, and price.
Run representative work
Include normal inputs, hard inputs, policy edges, and known failure cases from production.
Choose cost per accepted output
Cheapest per token or GPU-second is irrelevant if retries erase the saving.
Re-run on change
Model aliases, provider behavior, workflows, and prompts all change; regression data makes the change visible.