What SWE-bench Verified Means When a Coding Assistant Quotes a Score

A SWE-bench Score Is a Percentage of 500 Python Issues
A vendor claim of 70% on SWE-bench almost always means SWE-bench Verified. That set is 500 human-checked GitHub issues drawn from 12 popular Python repositories. The assistant reads an issue, edits the checked out repository, and counts as correct only when the project unit tests turn green.
That is the whole measurement. It ranks models against each other reasonably well, and it promises nothing about your language or your review queue.
Read the number as a lab result with a narrow scope. The rest of this guide covers what that scope includes and what it quietly leaves out.
Where the 500 Problems Actually Came From
- ● 2,294 issue and pull request pairs
- ● 12 popular Python repositories
- ● 500 kept after human validation
The original SWE-bench dataset collects 2,294 issue and pull request pairs from 12 popular Python repositories. The Hugging Face card puts 2,290 of them in the test split as of Aug 2026. Each task pins the repository to the commit just before the real fix landed.
The model receives the issue text and the repository. It never sees the human patch. Grading runs the project test suite twice, checking that the previously failing tests now pass and that the previously passing tests still do.
That design has one flaw that shows up quickly. Some issues in the original set were too vague to solve from the text alone, and some test suites punished patches that fixed the bug in a different but valid way.
SWE-bench Verified is the answer to that flaw. Its dataset card describes it as a subset of 500 samples from the SWE-bench test set that have been human-validated for quality, and the swebench.com changelog dates the release to 13 Aug 2024.
A Task Instance Is Mostly Two Lists of Test Names
Underneath the leaderboard, each task is a small record. The interesting part is not the issue text but the two lists of tests that decide the verdict.
{
"instance_id": "astropy__astropy-12907",
"repo": "astropy/astropy",
"base_commit": "d16bfe05a744909de4b27f5875fe0d4ed41ce607",
"problem_statement": "Modeling separability_matrix ...",
"FAIL_TO_PASS": ["tests/test_separable.py::test_nested"],
"PASS_TO_PASS": ["tests/test_separable.py::test_cstack"]
}
FAIL_TO_PASS holds the tests that must flip from red to green. PASS_TO_PASS holds the tests that must stay green, which is the check that stops a model from deleting an assertion to win.
A patch that satisfies both lists is scored as resolved. Nothing in that record asks whether the fix is idiomatic, whether it duplicates an existing helper, or whether a maintainer would merge it.
Four Datasets Share the Same Two Words
- ● Verified - 500 tasks
- ● Lite - 300 tasks
- ● Multimodal - 617 tasks
The phrase SWE-bench now covers several datasets of different sizes and difficulty. A percentage means little until you know which one produced it.
| Dataset | Task instances | Repositories | What it is used for | Source of the count |
|---|---|---|---|---|
| SWE-bench full | 2,294 pairs, 2,290 in the test split | 12 Python projects | The original benchmark, rarely quoted now | Hugging Face dataset card |
| SWE-bench Verified | 500 | 12 Python projects | Human-validated subset, the usual headline | Dataset card, released 13 Aug 2024 |
| SWE-bench Lite | 300 test, 23 dev | 11 Python projects | Cheap iteration and budget runs | Hugging Face dataset card |
| SWE-bench Multimodal | 617, with 510 in the test split | Interface projects such as Chart.js | Issues that carry screenshots and image assets | Hugging Face dataset card |
Counts come from the Hugging Face dataset cards and the SWE-bench site as of Aug 2026. Confirm the current figures at the source before you repeat them, since these datasets gain variants regularly.
Lite is the one to watch in a sales conversation. A Lite score sits on 300 easier tasks, so it tends to look better than a Verified score from the same system.
The Harness Around the Model Is Half the Result
A benchmark run is not just a model. It is a model wrapped in a scaffold that decides which files to show, how many attempts to allow, and whether failing test output flows back for another try.
Change the scaffold and the same weights produce a different percentage. That is why a model card, an independent leaderboard entry, and a partner blog post can disagree by ten points without anyone lying.
Four scaffold choices move the number most. How files reach the model, how many patches it drafts before one is picked, whether it may run tests, and how long each task may take.
None of those settings appear inside the headline percentage. They live in a methodology note, and sometimes in nothing at all.
Pass at 1 and Best of N Are Not the Same Claim
Pass at 1 means the system gets a single attempt per issue and is scored on that attempt. Best of N drafts several candidate patches and reports success if any of them passes.
The gap between those two protocols is wide on hard tasks. A generous retry budget makes a weaker model look competitive, which is exactly why the protocol belongs beside the number.
Ask which one the figure describes. If nobody can answer, the figure is marketing rather than measurement.
What the Benchmark Cannot See in Your Repository
Every Verified task lives in a public Python project with a working test suite, a clear issue thread, and no private dependencies. Your repository probably breaks at least two of those conditions.
Test coverage is the biggest gap. Grading assumes tests exist that will catch a wrong patch, and most internal codebases have thinner coverage than a popular open source library.
The benchmark also ignores everything after the tests go green. Readability, house conventions, and whether a reviewer can follow the diff carry no weight, though they decide whether a change actually ships. Our guide to reviewing AI generated code covers that second half.
Language coverage is the last blind spot. Verified and Lite are entirely Python, so a strong score is evidence about Python work and an assumption about anything else.
Contamination Is the Quiet Argument About These Scores
Every task comes from a public repository whose real fix merged years ago. Training corpora scraped from GitHub can contain those patches, the discussion around them, and the tests.
Nobody can prove a specific memorised answer from the outside. The honest position is that scores on public benchmarks drift upward for reasons that include genuine capability and reasons that do not.
This is one argument for private trials. A ticket from your own backlog last week cannot have been memorised, whatever the training set contained.
Which SWE-bench Figure Fits the Decision You Are Making
- ● Model shortlist - Verified pass at 1
- ● Cost sanity check - Lite
- ● Your codebase - your own trial
Shortlisting two or three models to trial: Use Verified with pass at 1. It is the closest thing to a common yardstick, and the gaps between leading systems are wide enough to guide a shortlist.
Comparing your own fine tuned or self hosted setup: Use Lite while iterating. The 300 task run is cheaper and faster, and relative movement is what matters during tuning.
Judging a tool for a non Python stack: Discount the score heavily. Look for language specific evidence instead, or run a trial on your own repository before committing.
Buying for a team with weak test coverage: Treat the benchmark as barely relevant. SWE-bench grades with strong test suites, and the assistant loses its safety net exactly where you need it most. Our note on measuring whether an assistant makes you faster suggests better local measurements.
Working on interface or front end code: Look for Multimodal results rather than Verified. Those 617 tasks include image assets, which is much closer to a bug report with a screenshot attached.
Five Questions That Turn a Score Into Evidence
Ask which dataset produced the figure, and expect Verified, Lite, or Multimodal by name. A vendor who answers only SWE-bench has told you nothing yet.
Ask for the protocol next. Pass at 1 with no retries is a strong claim, and best of ten is a different and weaker one.
Ask whether the run appears on the public leaderboard at swebench.com with submitted logs. Reproducible entries beat a percentage in a slide deck.
Ask what the scaffold did, particularly whether the agent could run tests and iterate. That single detail explains most disagreements between published numbers.
Finally, ask what the same setup scores on your stack. The honest answer is usually that nobody has measured it, which is the point at which a trial starts.
The Number Is a Filter, Not a Verdict
SWE-bench Verified is a genuinely useful benchmark. It uses real issues, real repositories, and a grader that cannot be charmed, and it separates strong systems from weak ones.
It is also 500 Python tasks with good tests and no organisational context. Comparing 65% against 70% across two different harnesses tells you less than a single afternoon spent on ten of your own tickets.
Use the published figure to pick which tools to try. Use your own backlog to pick which one to keep. Our current comparison of AI coding assistants is a reasonable place to build that shortlist.
FAQ
Which SWE-bench do vendors quote in their marketing?
Almost always SWE-bench Verified, the 500 sample human-validated subset. Ask which subset produced the figure, because Lite has 300 tasks and the full set has 2,294, and the three numbers are not interchangeable.
What counts as solving a SWE-bench task?
A resolved instance means the repository unit tests that failed before the patch now pass, and the tests that already passed still pass. It is a machine check on one Python project, not a judgement about readability or design.
Does a SWE-bench score cover languages other than Python?
No. Every task in Verified and Lite comes from Python repositories, so the score says nothing directly about Go, Rust, or TypeScript work. SWE-bench Multimodal adds interface projects with image assets, and it is scored separately.
Why do two published scores for the same model disagree?
The harness around the model changes the result. Retry counts, whether the agent sees test output, how files are retrieved, and the time limit per task all move the percentage without the underlying model changing at all.
Should a SWE-bench score decide which assistant we buy?
Treat it as a ranking signal between models, then run your own trial on ten real tickets from your own backlog. The benchmark predicts which models are worth trialling, not which one will suit your repository.
Sources
- SWE-bench — checked 2026-09-07
Some links may be affiliate links. We may earn a commission at no extra cost to you.
This article was written with AI assistance. It is researched and fact-checked, not based on personal hands-on testing unless explicitly stated.
Comments
Post a Comment