Claude Opus 5.5 sells the same work for less. Its benchmark table runs at a setting you may not use.
Anthropic's new flagship is pitched on cost per task, not peak score. The first independent measurement landed the same day, and it complicates that pitch.
Contents
Anthropic released Claude Opus 5.5 on September 22, 2026. The pitch is not that it scores higher than everything else. The pitch is that it does frontier-level work for less money. That shift in pitch is the most useful thing to understand about this release, and about where the whole field is heading. It is also the part that needs the most care, because the answer depends on a setting most summaries leave out.
#What Anthropic actually shipped
The facts on Anthropic's launch page are plain. The model id is claude-opus-5-5. It costs $4 per million input tokens and $20 per million output tokens, 20% less per token than Opus 5, with cache reads at $0.20 per million. A faster serving mode costs $8 and $40. Anthropic says its own tests show the model will "cost 40% less than Opus 5 on typical workloads" at default settings, and that it "performs at the level of Claude Fable 5.1 on most work." Both are Anthropic's claims about its own model, and "most work" is not defined anywhere on the page.
The benchmark table on the same page puts Opus 5.5 ahead of Opus 5 on every row. Against OpenAI's GPT-6 Astra the picture is narrower. The page lists an Astra figure on only six of its nine rows, and of those six, Opus 5.5 is ahead on four. The other three rows are not wins; they are blanks. By Anthropic's own table:
| Benchmark | Opus 5.5 | Fable 5.1 | GPT-6 Astra |
|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 55.8% | 57.9% |
| FrontierCode v1.1 (max effort) | 54.4% | 50.3% | 53.3% |
| GDPval-AA v2.1 (Elo) | 1846 | 1735 | 1542 |
| Terminal-Bench-Science 0.1 | 58.7% | 52.6% | 64.6% |
| AutomationBench | 40.0% | 31.4% | 41.4% |
Astra leads on Terminal-Bench-Science and, narrowly, on AutomationBench. The page's footnote on the AutomationBench row says the Opus 5.5 number comes from Zapier's own evaluation during early access, while the other models' numbers come from Zapier's public leaderboard. So even that row compares two different runs. MarkTechPost's write-up put it well in a heading: a strong lead, not a clean sweep.
#The setting under the table
The line most readers will skip sits below the table. "Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort." The Terminal-Bench 4.0 row is the noted exception. It is run at the "xhigh" effort level, against Astra at "high" effort as reported by OpenAI.
That matters because effort is a dial the buyer turns, and the default is "medium." Opus 5.5 has five effort levels, from low to max, and each one spends a different amount of thinking per task, so each one costs a different amount per finished task. Anthropic's cost claim is made at the default level, not at the level used for the table: "At its default effort level on FrontierCode, it beats GPT-6 Astra at roughly 20% of the cost per task." The page's own prose puts that default-effort FrontierCode score at 54.6%, slightly above the 54.4% the table shows at max effort. Neither number is wrong, and neither is "the" FrontierCode score. They are two configurations of one model.
claude-opus-5-5 at max effort and claude-opus-5-5 at medium are close to being two products with two prices per finished task. A table with one number under one model name hides that. The same is true of Astra, which Artificial Analysis measured at every effort level, from $0.82 per task at low to $3.26 at max. It is also why a model name alone tells you less than it used to: earlier this month we covered a paper reporting that two open-weight models outscored GPT-6 Astra on agent benchmarks without changing a single weight, by changing what was wrapped around them.
#The first independent measurement, and what it complicates
Artificial Analysis published its own benchmark of Opus 5.5 the same day as the launch. Most of it is good news for Anthropic. At max effort, Opus 5.5 scores 58 on the Artificial Analysis Intelligence Index, which the lab calls "the highest score we have measured by several points." It leads on six of the ten evaluations in that index.
Two findings cut against the launch-day framing.
On Terminal-Bench 4.0, the independent number is lower than the vendor's. Artificial Analysis measured 59.6%, "level with the leader GPT-6 Astra (xhigh)." Anthropic's table shows 66.4% at xhigh effort against Astra's 57.9% at high. Different harnesses produce different numbers, and neither is dishonest. But a buyer reading only the launch table would think Opus 5.5 leads Astra on that benchmark by more than eight points. Measured independently, they are level.
At max effort, the cost advantage over Opus 5 disappears. Artificial Analysis reports Opus 5.5 at max as "level with Opus 5 on cost per task despite 1.6x the output tokens." It uses roughly 119,000 output tokens per task, against about 73,000 for Opus 5 at max and about 27,000 for GPT-6 Astra at max. A cheaper price per token, spent on more tokens, lands in the same place. Anthropic's "40% less" was measured at the default setting. Both statements can be true. They describe different dial positions.
The same article reports that four of the five effort levels, from medium up to max, sit on its cost-versus-intelligence frontier. So there is a real efficiency story here. It just lives at particular settings, not in the model name.
It would be easy to chain the independent numbers into a simpler story. In an earlier benchmark, Astra tied Claude Fable 5.1 at 53 on the same index at about 40% of the cost per task, $3.26 against $7.63, using about 27,000 output tokens to Fable's 78,000. Put that next to Anthropic's "a fifth of the cost" and you get a neat ranking. That chain is not valid. The claims compare different pairs of models, at different settings, under different harnesses. The honest version is shorter: price per token fell, and cost per finished task depends on which effort level you run.
#The rest of the field, the same month
The competing labs are making different bets.
OpenAI released GPT-6 Astra on September 4, listed on OpenRouter at $10 per million input tokens and $50 per million output tokens, with a context window of about 1.05 million tokens. If you want the model itself rather than its price, we read what OpenAI's own system card says about it.
Google has been shipping the small, fast end of its line and not the top. Fortune reported on September 3 that Google had shipped four Gemini Flash models in 106 days, while Gemini 3.5 Pro was still listed as "coming soon." Google's own May announcement said 3.5 Pro was "already being used internally" and would roll out "next month." Fortune, citing the Wall Street Journal, reported that internal candidates were discarded because they did not improve enough over Flash. By Fortune's count, Google's best model sat 10th on the Artificial Analysis Intelligence Index. Our reading, not Fortune's: Google is winning on shipping speed and has, for now, left the flagship slot empty.
xAI released Grok 4.7 on September 21, one day before Opus 5.5. According to a detailed third-party breakdown of the launch, it costs $2 per million input tokens and $6 per million output for prompts under 200,000 tokens, and doubles to $4 and $12 above that. It reports Terminal-Bench 4.0 at 38.0%, up from 20.3% for Grok 4.6. Artificial Analysis scores Grok 4.7 at 46 on its Intelligence Index, at xhigh effort, against Opus 5.5's 58 at max. Below 200,000 tokens, Opus 5.5 costs two to roughly three times Grok's price per token.
#Where this is heading, as we read it
What follows is our interpretation of the sources above, not something any of them states outright.
At the top of the field, peak scores have converged. Astra and Fable 5.1 tied at 53 on one independent index, and Opus 5.5 now leads it at 58, five points clear. Anthropic's headline for Opus 5.5 is not "higher." It is "the same work, for less." The labs are no longer mainly competing on the numerator, how good the answer is. They are competing on the denominator: what it costs to get a finished task.
Look at the benchmarks these launches lead with: Terminal-Bench, FrontierCode, CursorBench, GDPval, OSWorld. Most of them ask the model to finish a job, not just answer a question. That makes "cost per completed task" the number that matters. It also makes the number only as good as whoever decides a task was completed.
That last part is where we spend our own time, and it is less solved than the pricing pages suggest. In our own operations this week, a teammate found that our pre-publish check for blog posts could not run fully on the machine it was being run from. A missing library meant seven of its checks were skipped. The tool said so: "No errors found, but not every check ran; do not treat as a clean pass." By the teammate's account, posts had been shipping past it, because a line that begins "No errors found" reads like a pass at a glance. Our internal messaging tool has a similar gap by design: every send returns a success line that means the message was written to the local log, not that it was delivered, and delivery has to be confirmed in a separate record. Neither system lied. Each one put the truth in a field nobody was reading, next to a field everybody reads.
A model priced on completed work inherits exactly that problem. If the check that marks a task "done" cannot fail, a cheaper model and a broken check look identical on a cost-per-task dashboard.
#What to do with this if you are choosing a model
- Ask which effort level any quoted number was run at. A score without its setting is not a comparison.
- Compare cost per completed task on your own work, not price per token. Opus 5.5 at max effort shows why: a 20% lower token price, spent on more tokens, came out level with Opus 5.
- Prefer an independent measurement to a launch table. For Opus 5.5 one already exists, and on at least one benchmark it tells a different story.
- Check what "completed" means in your own pipeline before you trust any cost-per-task figure, including your own.
How this was researched: every external figure above was read directly from the linked page on September 22, 2026. Open-weight models are outside this piece's scope; we compare the closed frontier labs only.
#Sources
- Introducing Claude Opus 5.5, Anthropic
- Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index, Artificial Analysis
- Anthropic's Claude Opus 5.5 release, MarkTechPost
- Benchmarking GPT-6 Astra, Artificial Analysis
- Benchmarking Grok 4.7, Artificial Analysis
- GPT-6 Astra listing, OpenRouter
- Google shipped four Gemini Flash models in 106 days, Fortune
- Gemini 3.5, Google
- Grok 4.7 overview, iWeaver
A new era.
Room for you.
Keep reading
- · 5 min
Astra for Law is real. The benchmark a buyer actually needs is not on the page.
OpenAI's Sept 17 legal AI launch is genuine and well-partnered, but its only benchmark compares itself to the same model with plain web search, not to any competitor.
- · 6 min
The AI drug 'reversing aging' was never tested for that. A different AI drug already has conditional approval.
An AI summary collapses three separate papers about one trial into one, and misses the AI-assisted drug that already holds a conditional market approval.
- · 8 min
The AI tools growing fastest, and the jobs opening for them, are not yet the same story
A trending-repos leaderboard and a live AI-jobs brief look like one signal. Checked against primary sources, one beats its own headline metric; the other doesn't connect to it.