- Opus 5.5 leads 7 of 9 benchmarks, with a 60% lower token price than GPT-6 Astra.
- For coding with Opus 5.5, stop at high effort. At max effort, Opus 5.5 passes 3% more tasks than at high, for 238% more cost per task.
- GPT-6 Astra's tokens cost 150% more than Opus 5.5's. Its only clear lead is reasoning and science: on Terminal-Bench-Science it passes 10% more tasks than Opus 5.5.
Best choice by task
Coding and data scripts
Opus 5.5high effort
56% on CursorBench 4.0$3.97 a task
- vs Fable 5.1 at max
- 8% fewer tasks, 335% more cost
- vs Opus 5.5 at max
- 3% more tasks, 238% more cost
At max effort, Opus 5.5 passes 3% more tasks than at high, for 238% more cost per task.
Cost data from the benchmark owner
Reports and client deliverables
Opus 5.5
1846 Elo on GDPval-AA 2.1No cost per task published
- vs Fable 5.1
- rated 111 points lower, 150% higher token price
- vs GPT-5.6 Sol
- rated 258 points lower, the same token price
GPT-6 Astra is rated 304 points lower on this benchmark.
Matched by a third party
Reporting and workflow automation
Opus 5.5max effort
42% on AutomationBench$1.44 a task
- vs GPT-6 Astra at max
- 3% fewer tasks, 20% more cost
- vs Opus 5.5 at xhigh
- 16% fewer tasks, 38% less cost
Zapier ran Opus 5.5 with fallbacks for refused steps. Anthropic's launch table, run without fallback, gives 40%.
Cost data from the benchmark owner
Research analysis
GPT-6 Astra
65% on Terminal-Bench-Science 0.1No cost per task published
- vs Opus 5.5
- 9% fewer tasks, 60% lower token price
For expert questions answered with tools, Opus 5.5 answers 3% more questions than Fable 5.1 on Humanity's Last Exam.
Vendor-reported
Working in tools without an API
Opus 5.5
82% on OSWorld 2.0No cost per task published
- vs Fable 5.1
- 1% fewer tasks, 150% higher token price
OpenAI's OSWorld 2.0 figures come from its launch, on a scoring mode that is not confirmed, and the two Chartography sources name different leaders.
Vendor-reported
Top models comparison by key benchmarks
Each bar shows how many more or fewer tasks the first model passes than the second, as a share of the second model's score. A gap inside the margin of error counts as a tie.
Launch table, 22 September 2026
Opus 5.5 ahead on 3 GPT-6 Astra ahead on 1 2 ties 3 without a comparison
Opus 5.5 against GPT-6 Astra: 60% lower token price; 17% less cost per task at max effort on AutomationBench.
Agentic coding
Knowledge work
Business workflows
Reasoning and science
Computer use and charts
Bars show the difference in tasks passed, up to 40% either side. The GDPval-AA row shows rating points, up to 400 either side.
Pass rate and cost per task
Price per token is an unfair comparison, because models use different numbers of tokens to complete a task. GPT-6 Astra's tokens cost 150% more than Opus 5.5's, yet at max effort on AutomationBench it costs 20% more per task.
Each chart plots the share of tasks a model completed correctly against what its provider charged per task, as published by the benchmark owner.
How we set model and effort in daily work is covered in our Claude Code workflow and Codex power-user workflow guides. Our guide to agentic RAG weighs capability against running cost for retrieval systems.
CursorBench 4.0
Cursor lists no GPT-6 Astra row, and publishes only the max-effort point for GPT-5.6 Sol. Source: Cursor, CursorBench 4.0.
Gain from each step up in effort
Opus 5.5
Fable 5.1
Opus 5
Show all 16 options as a table
| Model and effort | Pass rate | Cost per task |
|---|---|---|
| Opus 5.5 at max | 57.8% | $13.43 |
| Opus 5.5 at high sweet spot | 56.0% | $3.97 |
| Opus 5.5 at xhigh | 56.0% | $6.98 |
| Opus 5.5 at medium | 52.5% | $2.91 |
| Fable 5.1 at max | 51.8% | $17.28 |
| Fable 5.1 at xhigh | 51.6% | $13.01 |
| Fable 5.1 at high | 49.2% | $9.08 |
| Fable 5.1 at medium | 46.8% | $7.05 |
| Opus 5 at max | 46.6% | $11.95 |
| Opus 5 at xhigh | 46.1% | $11.43 |
| Fable 5.1 at low | 45.1% | $5.44 |
| Opus 5 at high | 44.7% | $9.00 |
| Opus 5.5 at low | 43.7% | $1.17 |
| Opus 5 at medium | 43.3% | $6.94 |
| GPT-5.6 Sol at max | 41.7% | $8.23 |
| Opus 5 at low | 40.7% | $4.87 |
AutomationBench
Zapier's Opus 5.5 rows use default fallbacks: a refused step is rerun on Anthropic's fallback model. For Fable 5.1, Opus 5 completed steps Fable's safety classifier refused, about 40% of tasks (260 of 657), and the cost covers Fable 5.1 only. Zapier publishes no low-effort points. Source: Zapier, AutomationBench v1.0.6.
Gain from each step up in effort
Opus 5.5
GPT-6 Astra
Show all 8 options as a table
| Model and effort | Pass rate | Cost per task |
|---|---|---|
| Opus 5.5 at max sweet spot | 42.5% | $1.44 |
| GPT-6 Astra at max | 41.4% | $1.73 |
| GPT-6 Astra at xhigh | 39.0% | $1.50 |
| GPT-6 Astra at high | 37.1% | $1.44 |
| Opus 5.5 at xhigh | 35.8% | $0.89 |
| GPT-6 Astra at medium | 34.1% | $1.27 |
| Opus 5.5 at high | 33.0% | $0.71 |
| Fable 5.1 at max | 31.4% | $2.45 |
All scores and sources
Anthropic's launch table of 22 September 2026. Each score shows its gap to the leader on that row. Scores for OpenAI models in this view are Anthropic's figures.
| Benchmark | Opus 5.5 Anthropic | Fable 5.1 Anthropic | Opus 5 Anthropic | GPT-6 Astra OpenAI | GPT-5.6 Sol OpenAI | Lead |
|---|---|---|---|---|---|---|
| Agentic coding | ||||||
| Terminal-Bench | Opus 5.5 +8.5 points, clear | |||||
| FrontierCode | Opus 5.5 +1.1 points, narrow | |||||
| CursorBench | No score found | Opus 5.5 +6.0 points, clear | ||||
| Knowledge work | ||||||
| GDPval-AA | Opus 5.5 +111 Elo, clear | |||||
| Business workflows | ||||||
| AutomationBench | GPT-6 Astra +1.4 points, narrow | |||||
| Reasoning and science | ||||||
| Humanity's Last Exam | No score found | Opus 5.5 +2.1 points, clear | ||||
| Terminal-Bench-Science | GPT-6 Astra +5.9 points, clear | |||||
| Computer use and charts | ||||||
| OSWorld | 72.6 OpenAI, other setting Not in Anthropic's table. OpenAI's launch table as reported by The Decoder reports 72.6. OpenAI's launch figure, reported in the press; scoring mode not confirmed. | 65.7 OpenAI, other setting Not in Anthropic's table. OpenAI's launch table as reported by The Decoder reports 65.7. OpenAI's launch figure, reported in the press; scoring mode not confirmed. | Opus 5.5 +1.1 points, narrow | |||
| Chartography | 71.0 Surge AI, other setting Not in Anthropic's table. Surge AI, Chartography leaderboard reports 71.0. Max effort, Surge's own protocol, tool use not stated. | 45.0 Surge AI, other setting Not in Anthropic's table. Surge AI, Chartography leaderboard reports 45.0. Max effort, Surge's own protocol, tool use not stated. | Opus 5.5 +0.6 points, narrow | |||
- An independent leaderboard reports the same score.
- Another source reports a different score, listed under conflicting figures.
- Reported by Anthropic. We could not confirm OpenAI's own figure.
- Not in Anthropic's table; another source publishes this score on a different setting, so it is not compared.
Benchmark definitions
- Terminal-Bench 4.0
- Whether an agent can finish real jobs in a command-line terminal: set up software, fix builds, process data, run a server. Share of tasks fully resolved, averaged over three runs.
- FrontierCode 1.1 Main
- Whether a code change written by the agent would be accepted into a real software project. Graded with the project's tests, maintainer rubrics and automated verifiers.
- CursorBench 4.0
- Coding requests taken from real sessions in the Cursor editor: vague briefs, several files, and some long tasks. Share of tasks solved.
- GDPval-AA 2.1
- Everyday professional deliverables such as reports, spreadsheets and plans, drawn from 44 occupations. An AI judge compares outputs blind, in pairs, and the results are scored as an Elo rating. Higher means preferred more often.
- AutomationBench 1.0.6
- Multi-step business workflows across connected apps: CRM updates, email, spreadsheets, ticketing. A task counts only when every check on the final state of the connected systems passes.
- Humanity's Last Exam with tools
- Expert-written questions across more than a hundred subjects, written to be hard for frontier models. Share of questions answered correctly.
- Terminal-Bench-Science 0.1
- Research tasks run end to end in a terminal: analysing data, running simulations, reproducing results. Share of tasks resolved.
- OSWorld 2.0
- Using a real desktop computer through the screen, mouse and keyboard to finish long tasks across ordinary apps. Anthropic reports a partial score, which credits progress on unfinished tasks, and a strict score for finished tasks only. This page uses partial.
- Chartography with tools
- Reading professional charts correctly and answering questions about them. Mean pass rate over 20 trials.
Conflicting figures
- Opus 5, FrontierCode 1.1 Main: 48.0 in the launch table; 53.4 in LLM Stats, FrontierCode 1.1 (self-reported figures).
- GPT-5.6 Sol, AutomationBench 1.0.6: 28.8 in the launch table; 19.6 in Anthropic, Claude Fable 5.1 launch page.
- Fable 5.1, Humanity's Last Exam: 65.6 in the launch table; 65.0 in Anthropic, Claude Fable 5.1 launch page.
- Fable 5.1, OSWorld 2.0: 80.7 in the launch table; 77.9 in Anthropic, Claude Fable 5.1 launch page.
- Opus 5, OSWorld 2.0: 74.0 in the launch table; 75.4 in Anthropic, Claude Fable 5.1 launch page.
- GPT-6 Astra, CursorBench 4.0: no score found. Cursor's own CursorBench 4.0 leaderboard has no GPT-6 Astra row.
Scores from other sources
- GPT-6 Astra, OSWorld 2.0: not in the launch table; 72.6 in OpenAI's launch table as reported by The Decoder (OpenAI's launch figure, reported in the press; scoring mode not confirmed).
- GPT-5.6 Sol, OSWorld 2.0: not in the launch table; 65.7 in OpenAI's launch table as reported by The Decoder (OpenAI's launch figure, reported in the press; scoring mode not confirmed).
- GPT-6 Astra, Chartography: not in the launch table; 71.0 in Surge AI, Chartography leaderboard (Max effort, Surge's own protocol, tool use not stated).
- GPT-5.6 Sol, Chartography: not in the launch table; 45.0 in Surge AI, Chartography leaderboard (Max effort, Surge's own protocol, tool use not stated).
Third-party checks
- GDPval-AA 2.1: Artificial Analysis reports the same 5 scores as the launch table.
Sources
- Anthropic, Claude Opus 5.5 launch page vendor
- LLM Stats, FrontierCode 1.1 (self-reported figures) aggregator
- Anthropic, Claude Fable 5.1 launch page vendor
- Vals AI, Terminal-Bench 4.0 leaderboard independent
- Artificial Analysis, GDPval-AA leaderboard independent
- Zapier, AutomationBench v1.0.6 independent
- Artificial Analysis, Humanity's Last Exam leaderboard independent
- Terminal-Bench, Terminal-Bench-Science 0.1 results independent
- Surge AI, Chartography leaderboard independent
- Cursor, CursorBench 4.0 independent
- OpenAI's launch table as reported by The Decoder secondary