Tool · Updated 24 September 2026

AI model benchmarks for Insights and Strategy

  • Opus 5.5 leads 7 of 9 benchmarks, with a 60% lower token price than GPT-6 Astra.
  • For coding with Opus 5.5, stop at high effort. At max effort, Opus 5.5 passes 3% more tasks than at high, for 238% more cost per task.
  • GPT-6 Astra's tokens cost 150% more than Opus 5.5's. Its only clear lead is reasoning and science: on Terminal-Bench-Science it passes 10% more tasks than Opus 5.5.

Best choice by task

Coding and data scripts

Opus 5.5high effort

56% on CursorBench 4.0$3.97 a task

vs Fable 5.1 at max
8% fewer tasks, 335% more cost
vs Opus 5.5 at max
3% more tasks, 238% more cost

At max effort, Opus 5.5 passes 3% more tasks than at high, for 238% more cost per task.

Cost data from the benchmark owner

Reports and client deliverables

Opus 5.5

1846 Elo on GDPval-AA 2.1No cost per task published

vs Fable 5.1
rated 111 points lower, 150% higher token price
vs GPT-5.6 Sol
rated 258 points lower, the same token price

GPT-6 Astra is rated 304 points lower on this benchmark.

Matched by a third party

Reporting and workflow automation

Opus 5.5max effort

42% on AutomationBench$1.44 a task

vs GPT-6 Astra at max
3% fewer tasks, 20% more cost
vs Opus 5.5 at xhigh
16% fewer tasks, 38% less cost

Zapier ran Opus 5.5 with fallbacks for refused steps. Anthropic's launch table, run without fallback, gives 40%.

Cost data from the benchmark owner

Research analysis

GPT-6 Astra

65% on Terminal-Bench-Science 0.1No cost per task published

vs Opus 5.5
9% fewer tasks, 60% lower token price

For expert questions answered with tools, Opus 5.5 answers 3% more questions than Fable 5.1 on Humanity's Last Exam.

Vendor-reported

Working in tools without an API

Opus 5.5

82% on OSWorld 2.0No cost per task published

vs Fable 5.1
1% fewer tasks, 150% higher token price

OpenAI's OSWorld 2.0 figures come from its launch, on a scoring mode that is not confirmed, and the two Chartography sources name different leaders.

Vendor-reported

Top models comparison by key benchmarks

Each bar shows how many more or fewer tasks the first model passes than the second, as a share of the second model's score. A gap inside the margin of error counts as a tie.

Launch table, 22 September 2026

Opus 5.5 ahead on 3 GPT-6 Astra ahead on 1 2 ties 3 without a comparison

Opus 5.5 against GPT-6 Astra: 60% lower token price; 17% less cost per task at max effort on AutomationBench.

Agentic coding

Terminal-Bench 4.0 15% more tasks66.4 vs 57.9
FrontierCode 1.1 Main Tie, 2% more tasks54.4 vs 53.3
CursorBench 4.0 No published score for GPT-6 Astra

Knowledge work

GDPval-AA 2.1 rated 304 points higher1846 vs 1542

Business workflows

AutomationBench 1.0.6 Tie, 3% fewer tasks40.0 vs 41.4

Reasoning and science

Humanity's Last Exam 18% more tasks67.7 vs 57.2
Terminal-Bench-Science 0.1 9% fewer tasks58.7 vs 64.6

Computer use and charts

OSWorld 2.0 No comparison on the same setting: GPT-6 Astra has a score only from another source
Chartography No comparison on the same setting: GPT-6 Astra has a score only from another source

Bars show the difference in tasks passed, up to 40% either side. The GDPval-AA row shows rating points, up to 400 either side.

Pass rate and cost per task

Price per token is an unfair comparison, because models use different numbers of tokens to complete a task. GPT-6 Astra's tokens cost 150% more than Opus 5.5's, yet at max effort on AutomationBench it costs 20% more per task.

Each chart plots the share of tasks a model completed correctly against what its provider charged per task, as published by the benchmark owner.

How we set model and effort in daily work is covered in our Claude Code workflow and Codex power-user workflow guides. Our guide to agentic RAG weighs capability against running cost for retrieval systems.

CursorBench 4.0

Best score for the cost A better option exists at the same price or less
CursorBench 4.0: pass rate against cost per task Opus 5.5 at low effort: 43.7% for $1.17 a task. Opus 5.5 at medium effort: 52.5% for $2.91 a task. Opus 5.5 at high effort: 56.0% for $3.97 a task. Opus 5.5 at xhigh effort: 56.0% for $6.98 a task. Opus 5.5 at max effort: 57.8% for $13.43 a task. Fable 5.1 at low effort: 45.1% for $5.44 a task. Fable 5.1 at medium effort: 46.8% for $7.05 a task. Fable 5.1 at high effort: 49.2% for $9.08 a task. Fable 5.1 at xhigh effort: 51.6% for $13.01 a task. Fable 5.1 at max effort: 51.8% for $17.28 a task. Opus 5 at low effort: 40.7% for $4.87 a task. Opus 5 at medium effort: 43.3% for $6.94 a task. Opus 5 at high effort: 44.7% for $9.00 a task. Opus 5 at xhigh effort: 46.1% for $11.43 a task. Opus 5 at max effort: 46.6% for $11.95 a task. GPT-5.6 Sol at max effort: 41.7% for $8.23 a task. 35% 40% 45% 50% 55% 60% $1 $2 $3 $5 $10 $20 Cost per task (US dollars, log scale) Pass rate low medium high xhigh max Opus 5.5 Fable 5.1 Opus 5 GPT-5.6 Sol

Cursor lists no GPT-6 Astra row, and publishes only the max-effort point for GPT-5.6 Sol. Source: Cursor, CursorBench 4.0.

Gain from each step up in effort

Opus 5.5

low → medium 20% more tasks 149% more cost
medium → high 7% more tasks 36% more cost
high → xhigh no gain 76% more cost
xhigh → max 3% more tasks 92% more cost

Fable 5.1

low → medium 4% more tasks 30% more cost
medium → high 5% more tasks 29% more cost
high → xhigh 5% more tasks 43% more cost
xhigh → max under 1% more tasks 33% more cost

Opus 5

low → medium 6% more tasks 43% more cost
medium → high 3% more tasks 30% more cost
high → xhigh 3% more tasks 27% more cost
xhigh → max 1% more tasks 5% more cost
Show all 16 options as a table
Model and effortPass rateCost per task
Opus 5.5 at max57.8%$13.43
Opus 5.5 at high sweet spot56.0%$3.97
Opus 5.5 at xhigh56.0%$6.98
Opus 5.5 at medium52.5%$2.91
Fable 5.1 at max51.8%$17.28
Fable 5.1 at xhigh51.6%$13.01
Fable 5.1 at high49.2%$9.08
Fable 5.1 at medium46.8%$7.05
Opus 5 at max46.6%$11.95
Opus 5 at xhigh46.1%$11.43
Fable 5.1 at low45.1%$5.44
Opus 5 at high44.7%$9.00
Opus 5.5 at low43.7%$1.17
Opus 5 at medium43.3%$6.94
GPT-5.6 Sol at max41.7%$8.23
Opus 5 at low40.7%$4.87

AutomationBench

Best score for the cost A better option exists at the same price or less
AutomationBench: pass rate against cost per task Opus 5.5 at high effort: 33.0% for $0.71 a task. Opus 5.5 at xhigh effort: 35.8% for $0.89 a task. Opus 5.5 at max effort: 42.5% for $1.44 a task. GPT-6 Astra at medium effort: 34.1% for $1.27 a task. GPT-6 Astra at high effort: 37.1% for $1.44 a task. GPT-6 Astra at xhigh effort: 39.0% for $1.50 a task. GPT-6 Astra at max effort: 41.4% for $1.73 a task. Fable 5.1 at max effort: 31.4% for $2.45 a task. 25% 30% 35% 40% 45% $1 $2 $3 Cost per task (US dollars, log scale) Pass rate high xhigh max Opus 5.5 medium high xhigh max GPT-6 Astra max Fable 5.1

Zapier's Opus 5.5 rows use default fallbacks: a refused step is rerun on Anthropic's fallback model. For Fable 5.1, Opus 5 completed steps Fable's safety classifier refused, about 40% of tasks (260 of 657), and the cost covers Fable 5.1 only. Zapier publishes no low-effort points. Source: Zapier, AutomationBench v1.0.6.

Gain from each step up in effort

Opus 5.5

high → xhigh 8% more tasks 25% more cost
xhigh → max 19% more tasks 62% more cost

GPT-6 Astra

medium → high 9% more tasks 13% more cost
high → xhigh 5% more tasks 4% more cost
xhigh → max 6% more tasks 15% more cost
Show all 8 options as a table
Model and effortPass rateCost per task
Opus 5.5 at max sweet spot42.5%$1.44
GPT-6 Astra at max41.4%$1.73
GPT-6 Astra at xhigh39.0%$1.50
GPT-6 Astra at high37.1%$1.44
Opus 5.5 at xhigh35.8%$0.89
GPT-6 Astra at medium34.1%$1.27
Opus 5.5 at high33.0%$0.71
Fable 5.1 at max31.4%$2.45

All scores and sources

Anthropic's launch table of 22 September 2026. Each score shows its gap to the leader on that row. Scores for OpenAI models in this view are Anthropic's figures.

Scores for five AI models on nine benchmarks. Percentages unless marked Elo.
Benchmark Opus 5.5 Anthropic Fable 5.1 Anthropic Opus 5 Anthropic GPT-6 Astra OpenAI GPT-5.6 Sol OpenAI Lead
Agentic coding
Terminal-Bench 4.0 Opus 5.5 +8.5 points, clear
FrontierCode 1.1 Main Opus 5.5 +1.1 points, narrow
CursorBench 4.0 No score found Opus 5.5 +6.0 points, clear
Knowledge work
GDPval-AA 2.1 · Elo Opus 5.5 +111 Elo, clear
Business workflows
AutomationBench 1.0.6 GPT-6 Astra +1.4 points, narrow
Reasoning and science
Humanity's Last Exam with tools No score found Opus 5.5 +2.1 points, clear
Terminal-Bench-Science 0.1 GPT-6 Astra +5.9 points, clear
Computer use and charts
OSWorld 2.0 72.6 OpenAI, other setting Not in Anthropic's table. OpenAI's launch table as reported by The Decoder reports 72.6. OpenAI's launch figure, reported in the press; scoring mode not confirmed. 65.7 OpenAI, other setting Not in Anthropic's table. OpenAI's launch table as reported by The Decoder reports 65.7. OpenAI's launch figure, reported in the press; scoring mode not confirmed. Opus 5.5 +1.1 points, narrow
Chartography with tools 71.0 Surge AI, other setting Not in Anthropic's table. Surge AI, Chartography leaderboard reports 71.0. Max effort, Surge's own protocol, tool use not stated. 45.0 Surge AI, other setting Not in Anthropic's table. Surge AI, Chartography leaderboard reports 45.0. Max effort, Surge's own protocol, tool use not stated. Opus 5.5 +0.6 points, narrow
  • An independent leaderboard reports the same score.
  • Another source reports a different score, listed under conflicting figures.
  • Reported by Anthropic. We could not confirm OpenAI's own figure.
  • Not in Anthropic's table; another source publishes this score on a different setting, so it is not compared.

Benchmark definitions

Terminal-Bench 4.0
Whether an agent can finish real jobs in a command-line terminal: set up software, fix builds, process data, run a server. Share of tasks fully resolved, averaged over three runs.
FrontierCode 1.1 Main
Whether a code change written by the agent would be accepted into a real software project. Graded with the project's tests, maintainer rubrics and automated verifiers.
CursorBench 4.0
Coding requests taken from real sessions in the Cursor editor: vague briefs, several files, and some long tasks. Share of tasks solved.
GDPval-AA 2.1
Everyday professional deliverables such as reports, spreadsheets and plans, drawn from 44 occupations. An AI judge compares outputs blind, in pairs, and the results are scored as an Elo rating. Higher means preferred more often.
AutomationBench 1.0.6
Multi-step business workflows across connected apps: CRM updates, email, spreadsheets, ticketing. A task counts only when every check on the final state of the connected systems passes.
Humanity's Last Exam with tools
Expert-written questions across more than a hundred subjects, written to be hard for frontier models. Share of questions answered correctly.
Terminal-Bench-Science 0.1
Research tasks run end to end in a terminal: analysing data, running simulations, reproducing results. Share of tasks resolved.
OSWorld 2.0
Using a real desktop computer through the screen, mouse and keyboard to finish long tasks across ordinary apps. Anthropic reports a partial score, which credits progress on unfinished tasks, and a strict score for finished tasks only. This page uses partial.
Chartography with tools
Reading professional charts correctly and answering questions about them. Mean pass rate over 20 trials.

Conflicting figures

Scores from other sources

  • GPT-6 Astra, OSWorld 2.0: not in the launch table; 72.6 in OpenAI's launch table as reported by The Decoder (OpenAI's launch figure, reported in the press; scoring mode not confirmed).
  • GPT-5.6 Sol, OSWorld 2.0: not in the launch table; 65.7 in OpenAI's launch table as reported by The Decoder (OpenAI's launch figure, reported in the press; scoring mode not confirmed).
  • GPT-6 Astra, Chartography: not in the launch table; 71.0 in Surge AI, Chartography leaderboard (Max effort, Surge's own protocol, tool use not stated).
  • GPT-5.6 Sol, Chartography: not in the launch table; 45.0 in Surge AI, Chartography leaderboard (Max effort, Surge's own protocol, tool use not stated).

Third-party checks

  • GDPval-AA 2.1: Artificial Analysis reports the same 5 scores as the launch table.