Artificial Analysis · Quality, price, and measured speed
Artificial Analysis
Independent model analysis spanning language intelligence, price and latency plus blind preference arenas for image and video generation.
What it measures
- •Language intelligence across agentic, coding, scientific-reasoning, and general tasks
- •Provider price, output speed, latency, and context-window comparisons
- •Blind human preference for text-to-image, image editing, and video modalities
Use it for
Building a broad shortlist while keeping quality, speed, and API price visible together.
Live leaderboards; read current values at the source instead of relying on a copied snapshot.
Read with care
Its Intelligence Index is a composite and the media arenas measure relative preference at standardized defaults. Neither substitutes for prompts, hardware, or failure modes from your workload.
From evidence to usage
Run the closest available path
These are availability mappings, not claims that an app.nz route reproduces the source’s score or harness.
Compare a frontier language panel
GPT · Claude · Gemini · DeepSeekSend one production prompt to four routed models and inspect output, latency, and billed cost side by side.
Run GPT Image 2
gpt-image-2Use the hosted image playground and OpenAI-compatible image endpoint for a prompt-level check.
Run open FLUX or SDXL
FLUX.1 · RealVisXL / SDXLTry hosted FLUX immediately or open a deployable RealVisXL Comfy workflow with mirrored weights.
Run open video workflows
Wan VACE · LTX-VideoMove from the video leaderboard to tested image-to-video, text-to-video, or video-to-video workflow JSON.
Make the final decision on your data
Keep public evidence as context, then score representative inputs with a frozen rubric. Inspect failures—not only the mean.