All eval sources

SWE-bench · Repository-level software repair

SWE-bench Verified

A human-validated 500-instance subset of real GitHub issues, evaluated by applying generated patches and running repository tests.

What it measures

  • Whether a generated patch resolves a real repository issue
  • End-to-end coding systems on containerized repository tasks
  • Language models under a shared minimal bash-agent scaffold

Use it for

Shortlisting coding models and agent systems for multi-file issue resolution rather than isolated code completion.

Use the leaderboard agent filter and compare like-for-like harness versions.

Read with care

Agent scaffold, tools, retries, repository selection, and release version materially affect the result. Model-only and full-system tables answer different questions.

From evidence to usage

Run the closest available path

These are availability mappings, not claims that an app.nz route reproduces the source’s score or harness.

Make the final decision on your data

Keep public evidence as context, then score representative inputs with a frozen rubric. Inspect failures—not only the mean.

Build an eval set