SWE-bench · Repository-level software repair
SWE-bench Verified
A human-validated 500-instance subset of real GitHub issues, evaluated by applying generated patches and running repository tests.
What it measures
- •Whether a generated patch resolves a real repository issue
- •End-to-end coding systems on containerized repository tasks
- •Language models under a shared minimal bash-agent scaffold
Use it for
Shortlisting coding models and agent systems for multi-file issue resolution rather than isolated code completion.
Use the leaderboard agent filter and compare like-for-like harness versions.
Read with care
Agent scaffold, tools, retries, repository selection, and release version materially affect the result. Model-only and full-system tables answer different questions.
From evidence to usage
Run the closest available path
These are availability mappings, not claims that an app.nz route reproduces the source’s score or harness.
Run a cloud coding agent
Provider-aware coding agentGive an app.nz agent a real repository issue and inspect its patch, tests, logs, and durable artifacts.
Compare coding models first
Frontier coding panelUse one issue and acceptance rubric to narrow the model panel before paying for full repository runs.
Make the final decision on your data
Keep public evidence as context, then score representative inputs with a frozen rubric. Inspect failures—not only the mean.