Skip to main content
Vantaige

SWE-bench Pro Explained: Why the Benchmark Changed (2026)

A
Aymen B
14 min read
SWE-bench Pro Explained: Why the Benchmark Changed (2026)

SWE-bench Pro Explained: Why the Benchmark Changed and What 64.3% Actually Means (2026)

Every AI coding agent now ships with a SWE-bench number on its pricing page, and through 2025 most of those numbers crept toward 80% on SWE-bench Verified. That climb stopped meaning what people thought it meant. Public benchmarks leak into training data, agents get tuned against the exact 500 tasks, and the headline percentage stops predicting how the tool behaves on your repository. SWE-bench Pro is the contamination-resistant replacement, and by April 2026 it is the number serious teams quote instead of Verified. The short version is in the TL;DR below.

TL;DR

  • SWE-bench Verified became gameable through training-data contamination.

  • SWE-bench Pro uses held-out tasks resistant to memorization.

  • Opus 4.7 leads SWE-bench Pro at a reported 64.3%.

  • Verified scores cluster near 80%; Pro spreads agents apart.

  • Read benchmarks as proxies, not guarantees, on your code.

What is SWE-bench Pro?

SWE-bench Pro is a contamination-resistant version of the SWE-bench coding benchmark, built so the evaluation tasks are unlikely to appear in a model's training data. It scores an AI agent on whether it can resolve real software issues, run the tests, and pass them, using problems held out from public datasets. The reported leader in April 2026 is Opus 4.7 at 64.3%.

The original SWE-bench, introduced by a Princeton-led team in the SWE-bench project, asks an agent to take a GitHub issue and the surrounding repository, produce a code patch, and have that patch pass the project's hidden test suite. It is a closer proxy for real engineering than multiple-choice coding quizzes because the agent has to navigate a codebase, not complete a single function.

SWE-bench Pro keeps that structure but fixes the leakage problem. The task pool is curated to reduce overlap with what frontier models have already memorized from public GitHub, and the harder, less-public problems are weighted so a model cannot coast on recall. The result is a lower absolute number for every agent and, more usefully, a wider spread between them.

Why did SWE-bench Verified become gameable?

Why did SWE-bench Verified become gameable?

SWE-bench Verified became gameable because its 500 tasks are public, fixed, and old enough that they leaked into model training corpora and agent-tuning loops. When the test set is known, "solving" it can mean recalling the upstream fix rather than reasoning about the code. Scores compressed near the top, so a 79% and a 77% told you almost nothing about real-world difference.

Three mechanisms drove the inflation. First, contamination: the patched commits for these GitHub issues are on the public internet, so any model trained after the issues were fixed may have seen the answer. Second, overfitting: vendors iterate agent scaffolding directly against the 500 tasks, tuning the test setup for that specific set. Third, saturation: once several agents sit in the high 70s, the metric loses resolution because the gaps fall inside the noise band.

None of this means SWE-bench Verified was dishonest. It was a strong benchmark that aged the way public benchmarks always age. The SWE-bench team itself documents the dataset's limitations on the official project site, and the move to a contamination-resistant variant is the standard fix once a leaderboard saturates.

What does 64.3% on SWE-bench Pro actually mean?

64.3% means that on the contamination-resistant task set, the reported leader (Opus 4.7) produced a patch that passed the hidden tests on roughly 64 of every 100 problems. It is lower than the same model's SWE-bench Verified figure because Pro strips out memorizable tasks. It is a relative ranking signal, not a promise that 64% of your tickets will close unattended.

Read the number three ways. As a ranking: a higher Pro score means the agent reasons better on unfamiliar code, which is what you actually buy. As a ceiling: a controlled benchmark with clean issues and existing tests is easier than a vague bug report against an undocumented service, so your real resolution rate will usually be lower. As a proxy: it measures patch-passes-tests, not code quality, maintainability, or whether the fix matches your architecture.

The honest framing for any agent score in 2026: treat it like a car's lab fuel-economy figure. It is measured consistently, it is useful for comparison, and you will not get that number in your own driving. Vantaige reports these as published benchmarks with that caveat attached every time.

How do the 2026 AI coding agents score?

By April 2026 the field consolidated to five agents in serious use: Claude Code, Cursor, Codex Desktop, Replit Agent 3, and Devin. On SWE-bench Verified the leaders cluster high (Claude Code 78.4%, Codex 71.0%, Cursor agent 67.2%), which is exactly the compression that made Verified less useful. SWE-bench Pro spreads them apart, with Opus 4.7 reported at 64.3%.

The table below lists the figures as reported by the SWE-bench project and each agent's official documentation. The token-efficiency column matters because two agents can reach a similar score while one burns several times the tokens to get there, which changes the real cost per resolved issue.

Agent / model

SWE-bench Verified (reported)

SWE-bench Pro (reported)

Token efficiency note

Claude Code (Opus 4.7)

78.4%

64.3% (Opus 4.7, Pro leader)

Reported ~5.5x fewer tokens than Cursor on identical high-complexity tasks

Codex Desktop

71.0%

Below Pro leader; vendor figure pending public confirmation

Mid-pack token use; varies with reasoning-effort setting

Cursor agent

67.2%

Below Pro leader; vendor figure pending public confirmation

Baseline for the ~5.5x comparison on high-complexity tasks

Replit Agent 3

Vendor-reported; confirm on official changelog

Not independently confirmed at publish

Optimized for app scaffolding, not minimal-diff repo patches

Devin

Vendor-reported; confirm on official changelog

Not independently confirmed at publish

Autonomous-session model; token use scales with task length

Two reading rules for this table. The Verified column is real but saturated, so a four-point gap there is inside the noise. The Pro column has fewer publicly confirmed entries because the benchmark is newer and vendors stage their submissions, so "pending" means not yet independently posted, not zero. When a vendor cites a Pro number, check it against the public leaderboard before quoting it.

If you are weighing specific models head to head rather than reading the leaderboard, the methodology in our DeepSeek V4 Pro vs Claude Opus 4.7 refactor benchmark shows how a single real repository test exposes differences a leaderboard hides.

Why does token efficiency matter as much as the score?

Token efficiency matters because the benchmark percentage tells you whether an agent can solve a problem, not what it costs to solve it. Claude Code is reported to use roughly 5.5x fewer tokens than Cursor on identical high-complexity tasks. At scale, a 5x token gap is the difference between a sustainable agent workflow and a bill that makes the workflow uneconomic.

Score and cost are separate axes and should be read together. An agent that resolves 64% of issues using one unit of tokens is operationally cheaper than one that resolves 60% using five units, even though their headline numbers look comparable. The leaderboard rarely surfaces this, which is why a benchmark-only decision can be the wrong decision.

This is also where routing strategy enters. Many teams send easy tickets to a cheap fast model and hard ones to the strongest agent. The trade-offs are mapped in our GPT-5.5 Instant vs Claude Opus 4.7 routing matrix, which treats cost-per-resolved-task as the metric that actually decides the architecture.

How is SWE-bench Pro different from SWE-bench Verified?

SWE-bench Pro differs from SWE-bench Verified in one decisive way: task selection is designed to resist contamination, so memorizing public fixes does not help. Verified is a human-validated 500-task subset of the original benchmark; Pro curates harder, less-public problems and weights them so recall cannot substitute for reasoning. Same scoring mechanic, harder and cleaner inputs.

The practical differences are best read side by side.

Dimension

SWE-bench Verified

SWE-bench Pro

Primary goal

Human-validated solvable subset

Contamination-resistant evaluation

Contamination exposure

High: fixed, public, well-known tasks

Low by design: held-out, harder problems

Score distribution (2026)

Compressed near 67-78% for top agents

Spread out; Opus 4.7 reported at 64.3%

Best used for

Historical trend, sanity check

Current ranking of agents on unfamiliar code

Failure mode

Overfitting and recall inflate scores

Newer, fewer publicly confirmed entries

Scoring mechanic

Patch must pass hidden tests

Patch must pass hidden tests (same)

Note what did not change: both benchmarks score "did the agent's patch make the real test suite pass." That is still a good proxy for engineering ability. What changed is the input quality. Pro is the better signal in 2026 not because the scoring is smarter but because the questions are no longer in the answer key.

What are the common mistakes when reading AI coding agent benchmarks?

What are the common mistakes when reading AI coding agent benchmarks?

The common mistakes are treating a benchmark as a guarantee, comparing scores across different benchmark versions, ignoring token cost, and assuming saturated leaderboards still have resolution. Each one leads teams to pick the wrong agent or to over-trust an autonomous workflow that needs review. Avoiding these four is most of what reading benchmarks well requires.

  • Treating the percentage as your hit rate. The fix: read it as a relative ranking and a ceiling. Your real resolution rate on messy internal tickets is almost always lower than the controlled-benchmark figure.

  • Comparing a Verified score to a Pro score. The fix: never put 78.4% (Verified) next to 64.3% (Pro) as if the second is a regression. Different task sets, different difficulty, not comparable.

  • Ignoring tokens. The fix: divide score by token cost. A reported ~5.5x token gap between Claude Code and Cursor on high-complexity tasks changes the economics more than a few percentage points of score.

  • Trusting a saturated leaderboard. The fix: when several agents sit within a few points on Verified, treat them as tied and decide on Pro, token cost, integration, and a test on your own repo instead.

  • Quoting an unverified Pro number. The fix: vendors stage submissions. If a Pro figure is not on the public SWE-bench leaderboard yet, label it vendor-reported until it is.

If you want a worked example of running your own controlled test instead of trusting a number, the editor benchmark methodology in Zed 1.0 vs Cursor vs VS Code real benchmark applies directly: fixed repo, fixed tasks, median of repeated runs.

Should you pick an AI coding agent based on its SWE-bench Pro score?

Use the SWE-bench Pro score as one input, not the decision. It is the best public signal for raw issue-resolution ability on unfamiliar code in 2026, but it does not measure token cost, integration with your stack, review burden, or code quality. The right process is: shortlist by Pro, then test the top two on your own repository before committing.

A defensible 2026 selection sequence. First, filter to agents with a credible Pro standing rather than a saturated Verified figure. Second, weigh token efficiency, because a reported ~5.5x gap dominates total cost at volume. Third, run a fixed set of your own real tickets through the final two and read resolution rate, diff quality, and review time, not just pass or fail. Fourth, re-check after any major model release, since the leaderboard moves.

Benchmarks rank capability; your repository ranks fit. The agent that tops SWE-bench Pro is the right starting hypothesis, and a one-day test on your codebase is what confirms or kills it. For self-hosted alternatives where you trade some score for control and cost, our Nous Hermes 4 self-hosted vs closed agents guide covers that trade-off in detail.

FAQ

Is SWE-bench Pro harder than SWE-bench Verified?

Yes. SWE-bench Pro is deliberately harder because it curates less-public, more complex problems and weights them so a model cannot rely on having memorized the upstream fix. Every agent scores lower on Pro than on Verified for that reason. Opus 4.7 is reported at 64.3% on Pro, below its Verified figure, and that drop is expected, not a regression. The difficulty is the point: a contamination-resistant benchmark only works if it is hard enough that recall does not substitute for reasoning.

What is the highest SWE-bench Pro score in 2026?

As of April 2026, Opus 4.7 (the model behind Claude Code) is the reported SWE-bench Pro leader at 64.3%. Treat that as the public ranking signal, not a guarantee of real-world performance, since controlled benchmark tasks are cleaner than internal tickets. Other agents have vendor-reported figures that are not yet independently confirmed on the public leaderboard. Always check a Pro number against the official SWE-bench project before quoting it, because vendors stage their submissions and numbers move with each model release.

Why is SWE-bench Verified still quoted if it is gameable?

SWE-bench Verified is still quoted because it has years of historical data, every agent has a figure for it, and it remains a useful sanity check and trend line. The problem is not that it is wrong; it is that contamination and overfitting compressed the top scores so a few points no longer signal real difference. Use Verified for the long-run trajectory of a model family and for catching obvious regressions, and use SWE-bench Pro for the current ranking of agents on code they have not memorized.

Does a high SWE-bench Pro score mean the agent works on my codebase?

Not directly. A high SWE-bench Pro score means the agent resolves unfamiliar issues well under controlled conditions: clean issue text, an existing test suite, a well-formed repository. Your codebase often has vague tickets, missing tests, and undocumented services, so your real resolution rate is usually lower than the benchmark figure. The score is the right way to shortlist agents, but a one-day test on your own real tickets, measuring diff quality and review time, is what tells you whether it works for you.

Why does Claude Code use fewer tokens than Cursor on the same tasks?

Claude Code is reported to use roughly 5.5x fewer tokens than Cursor on identical high-complexity tasks, which reflects differences in agent scaffolding: how much of the repository each tool reads into context, how it plans, and how it iterates toward a passing patch. The benchmark percentage does not capture this, but cost does. Two agents with similar scores can have very different cost-per-resolved-issue, so token efficiency belongs in the decision alongside the score, especially at production volume. Confirm current figures on each tool's official documentation.

Are benchmarks like SWE-bench Pro reliable for choosing tools?

They are reliable as a ranking input and unreliable as a guarantee. SWE-bench Pro is the most contamination-resistant public signal for coding-agent capability in 2026, which makes it good for shortlisting. It does not measure token cost, integration friction, code quality, or how an agent handles your specific stack. Treat any benchmark as a proxy: useful for comparison, measured consistently, but not the number you will see on your own work. Confirm the final choice with a controlled test on your repository.

Where can I see the official SWE-bench Pro leaderboard?

The authoritative source is the SWE-bench project itself, which publishes the leaderboard, the dataset, and the methodology at swebench.com. For agent-specific figures, cross-check against the vendor's official documentation: Anthropic's Claude Code page, Cursor, OpenAI Codex, Replit, and Devin. When the project leaderboard and a vendor page disagree, treat the project leaderboard as canonical and the vendor figure as vendor-reported.

Will SWE-bench Pro also become gameable over time?

Probably, eventually. Any public benchmark degrades once it is widely known: tasks leak into training data and vendors tune their scaffolding against the exact set. SWE-bench Pro buys resolution by being newer and harder, not by being permanently immune. The durable habit is not picking the perfect benchmark; it is reading whatever the current best benchmark is as a proxy, pairing it with token cost, and confirming with your own controlled test. When Pro saturates the way Verified did, the field will move again, and the same reading discipline will still apply.

References

  1. SWE-bench Project, "SWE-bench: leaderboard, dataset, and methodology", swebench.com

  2. Anthropic, "Claude Code", anthropic.com/claude-code

  3. Cursor, official site and changelog, cursor.com

  4. OpenAI, "Codex", openai.com/codex

  5. Replit, "Replit Agent", replit.com

  6. Cognition, "Devin", devin.ai

Get the best new AI tools and guides, weekly

One short email a week. The tools worth trying, the guides worth reading, nothing else.

No spam. Unsubscribe anytime.

A

Aymen B

Contributing writer at Vantaige, covering the AI tools ecosystem.