Model evals → working inference

A leaderboard is a lead,
not the decision.

Read independent evidence from Artificial Analysis and other maintained evals, open the exact models we host, then verify the shortlist on your prompts. One flow from public benchmark to production usage.

01

Find evidence

Use live sources, methodology, and uncertainty—not copied score snapshots.

02

Shortlist models

Match the benchmark modality and task distribution to the work you actually have.

03

Run the model

Open hosted gateway inference, a cloud agent, or a tested Comfy GPU workflow.

04

Keep the regression

Turn representative prompts and failures into a repeatable acceptance suite.

Independent evidence

Choose the question before the ranking

Every source below links to its live result and methodology. app.nz does not present old leaderboard values as current facts.

The useful eval loop

Public benchmarks are discovery data. Your acceptance suite is deployment data. Keep both, and record model version, provider route, prompt, settings, latency, cost, and failures.

Read the field guide
1

Start wide

Use multiple source types: objective tasks, human preference, agentic work, speed, and price.

2

Run representative work

Include normal inputs, hard inputs, policy edges, and known failure cases from production.

3

Choose cost per accepted output

Cheapest per token or GPU-second is irrelevant if retries erase the saving.

4

Re-run on change

Model aliases, provider behavior, workflows, and prompts all change; regression data makes the change visible.