All eval sources

Artificial Analysis · Quality, price, and measured speed

Artificial Analysis

Independent model analysis spanning language intelligence, price and latency plus blind preference arenas for image and video generation.

What it measures

  • Language intelligence across agentic, coding, scientific-reasoning, and general tasks
  • Provider price, output speed, latency, and context-window comparisons
  • Blind human preference for text-to-image, image editing, and video modalities

Use it for

Building a broad shortlist while keeping quality, speed, and API price visible together.

Live leaderboards; read current values at the source instead of relying on a copied snapshot.

Read with care

Its Intelligence Index is a composite and the media arenas measure relative preference at standardized defaults. Neither substitutes for prompts, hardware, or failure modes from your workload.

From evidence to usage

Run the closest available path

These are availability mappings, not claims that an app.nz route reproduces the source’s score or harness.

Make the final decision on your data

Keep public evidence as context, then score representative inputs with a frozen rubric. Inspect failures—not only the mean.

Build an eval set