Your Agent Benchmark Number Is Mostly Your Harness
TL;DR
Two number sets circulate under the name SWE-Bench Pro, about 20 points apart. Scale AI's standardised public leaderboard tops out at 61.50% (± 3.10) for Muse Spark 1.1 over 731 instances, where the top rows in fact ran the mini-swe-agent harness rather than the board's standard SWE-Agent scaffold. The aggregator board BenchLM carries 81.2% for Claude Fable 5.1 under the name "SWE-bench Pro" — higher than anything on the standardised board — with no scaffold, turn limit or subset named, and no source cited for where the number came from. Neither number is fabricated. They measure different systems, because one board controls the scaffold and the other does not.
Terminal-Bench 4.0 is a cleaner demonstration. One model, Claude Fable 5.1, carries three published scores: 55.8% from Anthropic's own announcement, 55.1% from Artificial Analysis, and 57.88% from an aggregator board. Anthropic's February 2026 study held the model, the harness and the task set fixed, varied only the container resources, and moved Terminal-Bench 2.0 by 6 percentage points at p below 0.01. Their published conclusion, hedged as what their data suggests and conditioned on resource methodology not yet being standardised: leaderboard differences below 3 percentage points deserve skepticism until the eval configuration is documented.
A leaderboard row is a property of four things — a model, a scaffold, a resource budget and a timeout — listed in decreasing order of attention and increasing order of effect.
SWE-Bench Pro Vendor-Reported vs Scale Standardized: The 20-Point Gap
Search for swe-bench pro vendor reported vs scale standardized and you get pages that quote one figure and move on. The honest answer is that two incompatible number sets are in circulation under one benchmark name, and they sit roughly 20 percentage points apart.
Scale AI runs a public SWE-Bench Pro leaderboard over 731 instances drawn from copyleft-licensed open-source repositories. The board's default configuration is SWE-Agent with uncapped cost and a 250-turn limit, but Scale asterisks the rows that departed from it — and that asterisk covers every row in the current top five, which ran the mini-swe-agent harness instead. These were the top rows when I read the board on 14 September 2026.
| Model | Score | Margin |
|---|---|---|
| Muse Spark 1.1 | 61.50% | ± 3.10 |
| gpt-5.4 (xHigh) | 59.10% | ± 3.56 |
| Muse Spark | 55.00% | ± 3.60 |
| claude-opus-4-6 (thinking) | 51.90% | ± 3.61 |
| gemini-3.1-pro (thinking) | 46.10% | ± 3.60 |
| claude-opus-4-5-20251101 | 45.89% | ± 3.60 |
| claude-4-5-Sonnet | 43.60% | ± 3.60 |
Source: the Scale SWE-Bench Pro public leaderboard↗, read 14 September 2026.
Against that, Anthropic reported Claude Fable 5.1 at 81.2% on "SWE-bench Pro", and aggregator boards relayed it as vendor-reported on 10 September 2026, naming neither the scaffold nor which of SWE-Bench Pro's three subsets — 731 public, 276 commercial or 858 held-out — the run covered. Same benchmark name. Nearly 20 points of daylight.
The gap is not a model gap. It is the scaffold. A benchmark like SWE-Bench Pro hands the system a repository and a failing test suite and asks for a patch. Everything between the model and the patch is the scaffold: how the agent lists files, how many turns it gets, whether it can run the tests before submitting, how a failed patch application is retried, how long it may run, and how much CPU and memory the container has. A vendor reporting its own number is running its own agent product end to end. Scale is running SWE-Agent for everyone, which makes the rows comparable to each other and not comparable to anything else.
Even "standardised" is not uniform
Scale asterisks the rows that ran mini-swe-agent under capped cost rather than the standard SWE-Agent scaffold at 250 turns with uncapped cost, and that asterisk covers the whole current top five — including the 61.50% figure I just quoted as the standardised ceiling. So the standardised board is itself a mix. Standardisation is a spectrum, not a binary.
The practical rule falls out of that: never put a vendor-reported figure and a standardised figure in the same table without a harness column. If your table has a model column and a score column and nothing else, the table is misleading whether or not every number in it is accurate.
What the margin of error already tells you
Look at the margins before the scores. Muse Spark 1.1 at 61.50 ± 3.10 and gpt-5.4 (xHigh) at 59.10 ± 3.56 overlap. Treating that as a ranking is reading noise as signal. The board publishes the uncertainty; most articles quoting the board drop it.
Three Sources, Three Terminal-Bench 4.0 Scores, One Model
Claude Fable 5.1's Terminal-Bench 4.0 score depends on who ran it.
| Reporter | Fable 5.1 on TB 4.0 |
|---|---|
| Anthropic (own announcement) | 55.8% |
| Artificial Analysis | 55.1% |
| Aggregator board (BenchLM) | 57.88% |
That is a 2.78-point spread on one model and one benchmark version across three organisations — and none of the three ran the same configuration. Artificial Analysis reports Claude Fable 5.1 at Xhigh effort with adaptive reasoning, pass@1 averaged over three repeats per task — run under mini-SWE-agent v2.4.6 with a 500-step cap and a 30-second command timeout, not under Claude Code, which is what the aggregator row used. The same scaffold difference the article discounts elsewhere is sitting inside this one model's three scores. The aggregator row is Claude Code at max effort. Anthropic published neither the harness nor the effort setting behind its 55.8%. Hold that number against Anthropic's own guidance, published seven months earlier, that gaps under 3 percentage points deserve skepticism until the configuration is documented. The disagreement between reporters about a single model is inside the band where Anthropic says you should not trust a model-to-model comparison.
Which means the configuration each reporter picked — documented by Artificial Analysis and by the aggregator, unpublished by Anthropic — is a larger variable than most of the model comparisons people build on top of these numbers.
One caveat I will not paper over: I could not read the canonical tbench.ai leaderboard tables directly — they did not render for me — so the per-row cost and token columns on the official board went unread, and the 57.88% figure is cited here as an aggregator figure, not as the canonical result.
The harness column explains more than the model column
Here are twelve of the eighteen rows on BenchLM's aggregated Terminal-Bench 4.0 board↗, snapshot of 11 September 2026, with the harness left in — the xhigh and high rows score identically and share a line below. I have cut six for length; none of them changes the shape.
| Model (harness, effort) | TB 4.0 |
|---|---|
| GPT-6 Astra (Codex, max) | 58.18% |
| Claude Fable 5.1 (Claude Code, max) | 57.88% |
| GPT-6 Astra (Codex, xhigh/high) | 57.88% |
| Claude Opus 5 (Claude Code, max) | 51.82% |
| Claude Fable 5 (Claude Code, max) | 44.55% |
| GLM-5.3 (Claude Code, max) | 41.82% |
| GPT-5.6 Sol (Codex, max) | 37.27% |
| Claude Opus 4.8 (Claude Code, max) | 23.64% |
| Grok 4.6 (Grok Build) | 20.30% |
| Gemini 3.8 Flash (mini-SWE-agent) | 19.09% |
| Claude Sonnet 5 (Claude Code, max) | 12.42% |
Gemini 3.8 Flash sits at 19.09%. It was run under mini-SWE-agent while the Claude rows ran under Claude Code and the GPT rows under Codex. That row is not comparable to the rows above it, and anyone quoting "Gemini 3.8 Flash scores 19% on Terminal-Bench" as a model fact is quoting a scaffold.
Notice also that GPT-6 Astra appears five times on that board, once per effort setting: 58.18% at max, 57.88% at xhigh and again at high, 54.24% at medium, 50.61% at low. That is a 7.57-point spread on one model from the effort dial alone — larger than the gap between any two adjacent rows in the board's top six. Effort is a harness decision, not a model property.
Infrastructure Configuration Alone Moves the Score by 6 Points
On 5 February 2026 Anthropic published the study that should be pinned to the top of every leaderboard. Six resource configurations on GKE. Same model. Same harness. Same task set. Only the container resources changed.
- Total lift at uncapped resources versus strict enforcement: +6 percentage points (p < 0.01) on Terminal-Bench 2.0.
- Infrastructure error rate: 5.8% at 1x spec enforcement, 2.1% at 3x (p < 0.001), 0.5% uncapped.
- Moving from 3x to uncapped dropped infra errors by a further 1.6 points while success jumped almost 4 points.
That last line is the interesting one. If generous resources only removed spurious failures, the success lift would track the infra-error drop. It does not. Success rose more than errors fell, which means the extra headroom let agents solve problems they otherwise could not — more parallel processes, more memory for a build, more room before the OOM killer arrives.
SWE-bench was far less sensitive to the same treatment: only 1.54 points higher at 5x RAM than at 1x, measured across 227 problems with 10 samples each.
The asymmetry is the lesson. A terminal or computer-use benchmark is a systems benchmark wearing a model benchmark's clothes. A patch-generation benchmark is closer to an actual model measurement. Read the infrastructure noise study↗ before you read another leaderboard.
Anthropic's own recommendation, verbatim: "Until resource methodology is standardized, our data suggests that leaderboard differences below 3 percentage points deserve skepticism until the eval configuration is documented and matched."
What Each Agentic Benchmark Actually Measures
Three benchmarks dominate agent marketing. They test different things and fail in different ways.
SWE-Bench Pro
731 instances from copyleft-licensed open-source repositories. The system is given a repository and must produce a patch that passes the hidden tests. On Scale's public board the default is SWE-Agent with a 250-turn limit and uncapped cost, but the asterisked rows — including the current top five — ran the mini-swe-agent harness instead, at the same uncapped cost and 250-turn limit.
What it measures: repository navigation, localisation of a defect, and patch generation under a turn budget. What it does not measure: anything about your codebase, your build system, or your test latency. The turn limit is a scaffold parameter, and a scaffold with 250 turns is a very different system from one with 40.
Terminal-Bench 4.0
66 tasks. Version 4.0 trimmed v3.0's 74 tasks down: 8 removed for being saturated, refusal-prone, publicly solved or platform-broken, and 19 fixed. Time, CPU and memory were raised benchmark-wide rather than per task, following Anthropic's infrastructure-noise methodology, and a flat 8-hour agent timeout was set. It is a major version rather than a 3.1 precisely because trials have to be re-run — the maintainers treat configuration changes as breaking changes, which is the correct instinct and rarer than it should be.
The maintainers cite Anthropic's infrastructure-noise methodology directly in the 4.0 announcement↗. Version 4.1 is planned with tamper-resistant verifiers, which tells you the maintainers expect agents to game the graders.
What it measures: end-to-end command-line competence inside a container. Which means it also measures the container, the timeout and the agent loop driving the shell.
tau2-bench
Multi-turn dialogue with tool calls, across banking-knowledge, retail, airline and telecom domains. tau2-bench's own leaderboard submission rules↗ make the point piecemeal rather than in one line: the same agent model and user simulator with identical arguments across all domains, one result per domain, every task run, at least four trials per domain with no task-id or task-count filters, the default scaffold and prompts (anything else is classed as a custom submission), the default base task split, and a stated pass^k level. Assemble them and you get the real requirement — domain, task release, agent model, user-simulator model, scaffold, prompts, trial count and the pass^k metric must all match before two scores are comparable.
That is eight dimensions. I have never seen a marketing page report more than two of them. If what you actually want is a vendor-level comparison rather than a benchmark ranking, the API-level differences are a better guide — I go through those in Claude API vs OpenAI.
A Saturated Benchmark Measures Nothing
The tau2-bench airline leaderboard on OpenRouter↗, refreshed 14 September 2026 at 09:00 UTC, had this at the top.
| Model | tau2-bench airline |
|---|---|
| Gemini 3.7 Flash | 80.6% |
| Claude Fable 5 | 80.2% |
| Claude Opus 5 | 79.6% |
| Amazon Nova Micro 1.0 | 78.7% |
| Qwen3.5 397B A17B | 78.5% |
| GLM 5.3 | 78.3% |
| Claude Fable 5.1 | 78.0% |
| DeepSeek V4 Pro | 77.7% |
| Gemini 3 Flash Preview | 77.3% |
Amazon Nova Micro 1.0 outranks Claude Fable 5.1. A micro model beats a flagship. This is an aggregator board rather than a primary source, and the scores use representative-run accuracy rather than an average across runs, which inflates the spread — so treat the exact ordering as directional. That caveat strengthens the point rather than weakening it.
The whole top nine sit inside 3.3 points. Against Anthropic's 3-point skepticism threshold, that is one undifferentiated blob with a ranking painted on it. When a benchmark saturates, its ordering stops carrying information about capability and starts carrying information about run-to-run variance.
Terminal-Bench 4.0's maintainers removed 8 tasks from v3.0 partly for saturation. That is the correct response to a benchmark that has stopped discriminating: retire the tasks, cut a major version, re-run everything. The wrong response is to keep quoting the ordering.
The Token Column Tells You More Than the Score Column
Terminal-Bench 4.0's maintainers published something more useful than the scores, and almost nobody picked it up: Claude Sonnet 5 consumed 21.6 billion tokens on its leaderboard run against 6.5 billion for Claude Opus 5. They also noted that Sonnet 5 struggled with output-token-exceeded errors.
On the aggregated board, Sonnet 5 under Claude Code at max effort scored 12.42% against 51.82% for Opus 5 under the same harness — those two scores come from the aggregator while the token totals come from tbench.ai's own announcement, so the pairing crosses two boards.
So the cheaper-per-million model used 3.3 times the tokens and returned about a quarter of the score. On a per-solved-task basis that is not a small difference, it is a different order of magnitude. There is also a feedback loop worth naming: a flat 8-hour timeout plus heavy token burn means the run is more likely to die before finishing, which drives the score down, which is then reported as a capability gap.
Dollars per million tokens is the wrong axis for any of this. I work through the arithmetic — the tokenizer change, the cache repricing, and how to instrument cost per completed task — in Tokens Per Task: The Real AI Cost Model in 2026. If you are budgeting an AI feature more broadly, the real cost of AI integration covers the parts that are not tokens at all.
Why the Harness Explains More of the Spread Than the Model
Put the measured harness effects next to each other.
| Variable changed | Measured effect | Source |
|---|---|---|
| Container resources, strict to uncapped | +6.0 points on Terminal-Bench 2.0 (p < 0.01) | Anthropic, 5 Feb 2026 |
| Reporter plus harness, same model and benchmark | 2.78 points on TB 4.0 (55.1% under mini-SWE-agent at Xhigh, to 57.88% under Claude Code at max) | Three published figures |
| Effort setting, max vs low | 7.57 points on GPT-6 Astra (58.18% to 50.61%) | Aggregated TB 4.0 board |
| Scaffold, mini-SWE-agent vs Claude Code / Codex | Not isolated, but large enough to put a Flash model at 19.09% | Aggregated TB 4.0 board |
| RAM, 1x to 5x on SWE-bench | +1.54 points across 227 problems | Anthropic, 5 Feb 2026 |
Now compare those to the model gaps people argue about. The difference between the first and second row on Scale's SWE-Bench Pro board is 2.4 points with margins of ± 3.10 and ± 3.56. The harness effects are larger than the model effects, and the harness effects are the ones nobody documents.
Prithvi Rajasekaran of Anthropic Labs put the same idea more sharply in Harness design for long-running application development↗, published 24 March 2026: every component in a harness encodes an assumption about what the model cannot do on its own, and those assumptions are worth stress testing. Their measured example makes it concrete. The same brief, given to Claude Opus 4.5 alone, took 20 minutes and $9 and produced broken core functionality. Run through the full generator and evaluator harness on that same model, it took 6 hours and $200, working from a planner-expanded 16-feature spec across ten sprints, and delivered a game maker whose output he could actually play. Same model. The delta is entirely harness.
That is also why harness complexity should shrink as models improve. Rajasekaran removed the sprint construct that Opus 4.5 needed once Opus 4.6 stopped needing it. A harness component that has outlived its assumption is pure cost. I go deeper on this in Multi-Agent Systems Are a Context Decision, Not an Org Chart, and on what to do with the context window itself in Context Engineering: Why Compaction Is the Wrong First Lever.
Build the Eval for Your Own System Instead
None of the public numbers predict how your agent will do on your codebase, your tools and your users. Build a small eval and use it. Anthropic's published guidance↗ from 9 January 2026 is that 20 to 50 simple tasks drawn from real failures is a good start, because early-stage effect sizes are large enough that you do not need statistical power you cannot afford.
Take the tasks from your incident log
Not from a benchmark, not from your imagination. Real failures, transcribed. The tasks you already know your system gets wrong are the ones where a change will show a measurable delta. Once your eval stops catching regressions, add the newest failures and retire the ones that always pass.
Decide between pass@k and pass^k before you run anything
These measure different products.
- pass@k is the probability that at least one of k attempts succeeds. That is the right metric for a tool a developer supervises and can re-run.
- pass^k is the probability that all k trials succeed. That is the right metric for anything customer-facing, where one failure in ten is a support ticket.
They diverge sharply as k grows. A system at 90% per-attempt success is 99.999% on pass@5 and 59% on pass^5. Same model, same tasks, two numbers that support opposite decisions. Publish which one you used, internally as well as externally.
Grade outputs, not process paths
Anthropic's guidance is explicit about avoiding rigid step-sequencing. If your grader requires the agent to call tools in a fixed order, you are measuring conformance to your assumptions rather than task success, and you will penalise the model every time it finds a shorter route. Grade the artifact.
Give an LLM judge an explicit "Unknown"
A judge forced to pick pass or fail will invent a verdict on the cases it cannot evaluate. Give it a third option. Use structured rubrics that score isolated dimensions rather than one holistic number. Then read transcripts regularly — not to grade the agent, but to grade the grader. The judge drifts, and nothing in the pipeline will tell you.
The grading prompt is a production prompt and deserves the same discipline as the rest of them, which I cover in prompt engineering for production apps.
Pin the configuration or the number means nothing
This is the whole lesson from the three Terminal-Bench scores. Every eval run should emit a manifest alongside its score.
eval_run:
model: claude-opus-5
effort: high
harness: internal-agent-loop@4.2.1
max_turns: 250
wall_clock_timeout_s: 28800
container: 4vCPU / 16GiB RAM / 20GiB disk
resource_enforcement: uncapped
tool_count: 12
trials_per_task: 5
metric: pass^5
task_count: 34
task_source: incident-log-2026-Q2
tokens_in: 4183220
tokens_out: 511903If you cannot fill in every field for two runs, those two runs are not comparable. If you can, you have something the public leaderboards do not: a number that means one specific thing.
Log tokens_in and tokens_out per run whether or not you think you care about cost. The token column is how you catch a change that improved the score by grinding through four times the work. The Sonnet 5 row on Terminal-Bench 4.0 is that failure mode, published.
Do Models Behave Differently When They Know They Are Being Tested?
There is a further problem I could not source to a primary publication, so take it as an open question rather than a finding: whether models behave differently when they detect they are being evaluated. If they do, it is a direct threat to any eval built out of synthetic, benchmark-shaped tasks.
The mitigation is the same as the advice above, for a different reason: pull tasks from real traffic and keep the surface shape of the real system. If your eval prompt says "you are being evaluated on the following task", you are measuring evaluation behaviour.
Related to this, Anthropic's eval guidance↗ also warns about eval saturation on your own suite. When every task passes, the suite has stopped being an instrument and become a ritual.
A Checklist for Reading Any Leaderboard Row
Before you quote a number, answer these. If you cannot answer more than half, do not quote it as a model property.
- Which scaffold ran it? Claude Code, Codex, SWE-Agent, mini-SWE-agent and a bespoke internal loop are five different systems.
- Which effort or thinking setting? GPT-6 Astra spans 7.57 points across its five effort settings on one board, 58.18% at max down to 50.61% at low — and the max-versus-xhigh ordering reverses between that board and Artificial Analysis, where xhigh sits at 59.6% against max at 59.1%.
- Turn limit and wall-clock timeout? Terminal-Bench 4.0 uses a flat 8-hour agent timeout. SWE-Agent on Scale's board uses 250 turns.
- Container resources, and were they enforced? Anthropic measured 6 points between strict enforcement and uncapped.
- Trials per task, and is the score a representative run or a mean across runs? The tau2-bench airline board uses representative-run accuracy.
- pass@k or pass^k? If unstated, assume the flattering one.
- Who published it — the vendor, an independent lab, or an aggregator copying both? One model had three Terminal-Bench 4.0 scores across three publishers.
- Is the gap you care about larger than 3 percentage points? Below that, Anthropic says to be skeptical until the configuration is documented, and Anthropic builds one of the models on the board.
- What did the run cost in tokens? 21.6 billion versus 6.5 billion changed the meaning of two scores on the same board.
- Is the benchmark saturated? Nine models inside 3.3 points on tau2-bench airline is not a ranking.
That checklist is also a specification for what to publish when you report your own numbers. If you are choosing a model for a production agent and want the evaluation done properly against your workload rather than against a public board, let's talk about your project.
Key Takeaways
- SWE-Bench Pro has two incompatible number sets. Scale's standardised public set tops out at 61.50% (± 3.10) over 731 instances, with its top five rows on the mini-swe-agent harness rather than the board's standard SWE-Agent scaffold; an aggregator board carries 81.2% for Claude Fable 5.1 under the same benchmark name, with no harness, subset or source stated. Never show both without a harness column.
- One model had three Terminal-Bench 4.0 scores. Claude Fable 5.1 at 55.8% (Anthropic), 55.1% (Artificial Analysis) and 57.88% (aggregator) — a 2.78-point spread across three configurations nobody aligned — mini-SWE-agent at Xhigh effort over three repeats at Artificial Analysis, Claude Code at max effort on the aggregator, and an unpublished harness at Anthropic.
- Resources alone move a score 6 points. Anthropic's February 2026 study held model, harness and tasks fixed and measured +6 percentage points (p < 0.01) on Terminal-Bench 2.0 between strict enforcement and uncapped, with infra errors falling from 5.8% to 0.5%.
- Read the harness column first. Gemini 3.8 Flash's 19.09% on Terminal-Bench 4.0 came from a mini-SWE-agent run while the rows above it used Claude Code and Codex.
- A saturated benchmark is not a ranking. Amazon Nova Micro 1.0 at 78.7% outranks Claude Fable 5.1 at 78.0% on tau2-bench airline, with nine models inside 3.3 points.
- The token column outranks the score column. Sonnet 5 burned 21.6B tokens to Opus 5's 6.5B on Terminal-Bench 4.0 and scored 12.42% against 51.82%.
- Your eval beats their leaderboard. 20 to 50 tasks from real failures, a chosen pass@k or pass^k, output grading, an "Unknown" option for the judge, and a pinned configuration manifest on every run.
About the Author
I'm Uvin Vindula — a Web3 and AI engineer based between Sri Lanka and the UK. I build production agent systems, and the published evidence in this article — six points from container resources alone, 7.57 points from one effort dial — is why I no longer quote leaderboard figures without the configuration attached. You can see my work at iamuvin.com or reach out about a project at hello@iamuvin.com↗.
If you are choosing between models for an agent product and want an eval built on your own failure log rather than a public board, let's talk about your project.
Working on a Web3 or AI project?
More in AI & Machine Learning
- Tokens Per Task: The Real AI Cost Model in 2026
- Multi-Agent Systems Are a Context Decision, Not an Org Chart
- Context Engineering: Why Compaction Is the Wrong First Lever
- MCP Went Stateless: Migrating Your Server

Uvin Vindula
Web3 and AI engineer based in Sri Lanka and the UK. Author of The Rise of Bitcoin. Founder of ASI Research Labs. Director of Blockchain and Software Solutions at Terra Labz. Founder of uvin.lk — Sri Lanka's Bitcoin education platform with 10,000+ learners.