LiveBench · Fresh, objective language tasks
LiveBench
A frequently refreshed benchmark designed around recent questions, objective ground-truth scoring, and reduced test contamination.
What it measures
- •Math, coding, reasoning, language, instruction following, and data analysis
- •Objective answers rather than a model judge where possible
- •Performance on questions sourced or transformed from recent material
Use it for
Checking whether a language-model shortlist still generalizes beyond old, heavily circulated benchmark sets.
Questions and releases evolve; compare results from compatible benchmark versions.
Read with care
Freshness reduces one contamination risk; it does not make the category mix identical to your users or account for end-to-end agent scaffolding.
From evidence to usage
Run the closest available path
These are availability mappings, not claims that an app.nz route reproduces the source’s score or harness.
Test the same models on your prompt
Reasoning and fast routesUse benchmark evidence to pick candidates, then measure the exact task, latency, and cost you will ship.
Run a repeatable dataset eval
Any gateway modelsTurn representative inputs into a scored regression set with an explicit judge rubric.
Make the final decision on your data
Keep public evidence as context, then score representative inputs with a frozen rubric. Inspect failures—not only the mean.