All eval sources

LiveBench · Fresh, objective language tasks

LiveBench

A frequently refreshed benchmark designed around recent questions, objective ground-truth scoring, and reduced test contamination.

What it measures

  • Math, coding, reasoning, language, instruction following, and data analysis
  • Objective answers rather than a model judge where possible
  • Performance on questions sourced or transformed from recent material

Use it for

Checking whether a language-model shortlist still generalizes beyond old, heavily circulated benchmark sets.

Questions and releases evolve; compare results from compatible benchmark versions.

Read with care

Freshness reduces one contamination risk; it does not make the category mix identical to your users or account for end-to-end agent scaffolding.

From evidence to usage

Run the closest available path

These are availability mappings, not claims that an app.nz route reproduces the source’s score or harness.

Make the final decision on your data

Keep public evidence as context, then score representative inputs with a frozen rubric. Inspect failures—not only the mean.

Build an eval set