Articles

Claude Opus 5: Benchmarks, Pricing, and 1M Context

Anthropic launched Claude Opus 5 on July 24 with a 1M-token context, a new xhigh effort mode, and unchanged $5/$25 pricing. Benchmarks, specs, and what changed.

Chisato Chisato · · 6 min read
A glowing AI model badge representing a newly released large language model

Anthropic has a new flagship. On Thursday, July 24, 2026, the company released Claude Opus 5, the newest and most capable model in its Opus line, positioning it as a frontier-class system for agentic coding and computer use while holding the price of its predecessor flat. The headline the company chose to lead with was not a benchmark number but a claim about economics: near-flagship performance at the same $5 and $25 per million tokens Opus buyers were already paying.

The launch lands in the middle of the busiest stretch of frontier releases the industry has seen. It arrives two weeks after SpaceXAI’s Grok 4.5 shipped as a self-described “Opus-class” model at a fraction of Opus pricing, and while OpenAI is still rolling out its GPT-5.6 family. Anthropic’s answer is to raise the ceiling on its own most-used tier rather than launch a separate premium line — a notable choice a month after it introduced Claude Fable 5 as its top-end flagship.

What Anthropic shipped

Claude Opus 5 is a reasoning model with extended thinking on by default, the test-time-compute approach that defines the current generation and that we cover in our primer on reasoning models. The specifications read like a checklist of 2026 frontier features:

  • A 1-million-token context window, up from the 200K that defined earlier Opus releases and matching the long-context tier now standard across the field. Our explainer on what a context window is covers why that number matters for agents that hold large codebases in working memory.
  • A 128,000-token maximum output, enough for the model to produce large multi-file changes or long structured documents in a single pass.
  • A new “xhigh” effort setting that sits above the previous maximum, giving developers a dial to trade latency and cost for accuracy on the hardest problems.

Anthropic is also making Opus 5 the new default model on Claude Max and the strongest option available on Claude Pro, folding the release directly into its consumer subscriptions rather than gating it behind the API alone.

The pricing story

The most deliberate part of the announcement is what did not change. Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens in standard mode — identical to Claude Opus 4.8 — with a faster priority tier at $10 and $50. Anthropic’s framing is that a full generational jump in capability arrives at half the price of Fable 5 while leaving the Opus price card untouched.

That positioning is aimed squarely at a market that has spent 2026 watching per-token prices fall. Chinese open-weight labs and aggressive challengers have pushed capable models to a few dollars per million tokens, and Anthropic’s pitch is that it can hold a premium price only if each new Opus delivers materially more work per dollar. Holding the number flat while roughly doubling benchmark scores is the company’s attempt to reset the value equation without touching the sticker.

The benchmarks

On the numbers Anthropic released, Opus 5 is a clear step up from Opus 4.8 and competitive with the best models in the field.

On SWE-bench Verified, the most-watched software-engineering benchmark, Anthropic reported Opus 5 at 96.0%, edging the roughly 95% that Fable 5 posts and placing it at the top of the coding leaderboard. On the harder SWE-bench Pro, the company reported 79.2%. These are the scores that matter most to Anthropic’s core audience: the model is tuned to close real bugs and land real pull requests, the workload we survey in our look at the state of AI coding assistants.

The more striking result is on agentic, tool-using tasks. On FrontierBench v0.1 — a 74-task successor to the Terminal-Bench series that measures how well a model drives a real terminal and computer over long horizons — Anthropic reported 43.3% at maximum effort, ahead of GPT-5.6 Sol at 37.5%. But the number the company emphasized was the new xhigh setting, which reached 44.4% mean reward — its best result on the benchmark — while spending about 25% fewer output tokens than max effort. In other words, the highest-quality mode was also cheaper to run than the previous ceiling, a token-efficiency claim that, if it holds in production, matters as much to buyers as the raw score.

Anthropic also highlighted ARC-AGI 3, an abstract-reasoning benchmark designed to resist memorization, where it said Opus 5 scored roughly three times as high as the next-best model, and CursorBench 3.2, an editor-integrated coding evaluation, where it reported Opus 5 landing within about 0.5% of Fable 5’s peak at half the cost per task. As always with vendor-reported figures, these come from Anthropic’s own testing; independent replications will follow in the days ahead, and the field’s evaluation practices are worth reading skeptically, as we discuss in what an LLM eval is.

Where it fits in the lineup

Opus 5 sharpens a lineup that had become genuinely confusing. Anthropic spent early 2026 with Opus 4.8 as its workhorse before introducing Fable 5 as a separate top-end flagship in June — a move that left many developers asking where a “Sonnet 5” or a next Opus fit, a question we untangled in our 2026 Claude lineup guide. Thursday’s release answers part of it: Opus is back as the model most users will actually run, with Fable 5 reserved as the maximum-capability, maximum-cost option and Opus 5 covering the broad middle where cost and quality both matter.

The positioning is explicitly agent-first. Anthropic pitched Opus 5 not as a chat model but as the engine for autonomous coding and computer-use workflows — the kind of long-horizon, tool-calling behavior described in our overview of what an AI agent is. That focus tracks where enterprise spending has moved: the money is increasingly in models that can operate software over many steps without a human in the loop on every action.

What it means

Claude Opus 5 is best read as a defensive price move dressed as a capability launch. The capability gains are real — a top SWE-bench score and a token-efficient xhigh mode are meaningful — but the strategic core is holding Opus pricing flat while roughly doubling performance. In a year when Grok 4.5 and Chinese open-weight models have hammered the cost of “good enough” intelligence, Anthropic is betting it can defend a premium tier only by making each dollar buy more finished work, not by discounting.

Who wins: developers and agent builders, who get a stronger default without a price hike, and Anthropic’s own subscription business, which now ships its best coding model straight into Max and Pro. The xhigh-cheaper-than-max claim, if it survives independent testing, is the sleeper win — it attacks the biggest objection to reasoning models, which is that maximum quality means runaway token bills.

Who feels the pressure: rivals selling on price. If Opus 5 genuinely delivers near-Fable performance at $5/$25, the “Opus-class at a fraction of the cost” pitch that challengers have leaned on gets harder to make. And OpenAI, still mid-rollout on GPT-5.6, now faces a competitor claiming the top of the most commercially important benchmark — agentic software engineering — on the same week its own models were making headlines for the wrong reasons.

What to watch next: independent benchmark replications, especially on FrontierBench and the xhigh token-efficiency claim; whether Anthropic’s coding lead translates into net-new API share or mostly upgrades existing Opus users in place; and pricing responses from OpenAI and the open-weight camp. The frontier is now a monthly release cadence, and the interval between “new state of the art” announcements is measured in weeks, not quarters.

Chisato Chisato · · 3 min read

Is There a Claude Sonnet 5? Anthropic's 2026 Lineup

Looking for Claude Sonnet 5? Here's the honest answer — plus a clear map of Anthropic's 2026 models: Haiku 4.5, Sonnet 4.6, Opus 4.8, and the new Fable 5.

#AI #Claude #Anthropic
Chisato Chisato · · 4 min read

Claude Fable 5: Anthropic's Most Capable Model Yet

Anthropic's Claude Fable 5 is its most capable model yet, built for long-horizon, autonomous agent work. Here's what's new, what it costs, and when to use it.

#AI #Claude #Anthropic