SpaceRock

Opus 5: Anthropic’s New Flagship Explained

By Veer Solanki · · 1346 words

Topics: AI, AI Alignment, AI Benchmarking, AI Guardrails, AI industry, AI Infrastructure, AI Models, AI Pricing

Opus 5: Anthropic’s New Flagship Explained

Friday afternoon, Anthropic launched Opus 5. 5$ per million input tokens, 25$ for output. Same price as Opus 4.8. Not nearly as good as the selfsame Opus 4.8. This is the sentence they led the press release with: a model coming close to the frontier intelligence of Claude Fable 5 at half the price. They’re not saying they found a new frontier; they’re saying they got most of the way to their own best model for half the cost their own best model charges.

That’s not a line companies tend to lead with when they have other options.

Six weeks ago I wrote about Fable 5 being quietly removed from the marketplace three days after its launch, and about who gets to decide what gets shipped in 2026. This is another story about control, but one with less drama and more immediate value to the person paying for the tokens.

The dial

The feature Anthropic chose to spotlight is an effort setting. Low, medium, high, xhigh, max; you pick how much compute the model will burn before answering a question. Every performance chart in the release is a curve, not a bar. Cost per task on one axis, performance on the other.

I want to linger on that because it represents a shift in the conversation around these models. For two years, the charts were bars, taller ones winning the argument. Now the chart is a curve, and the question is always about where it bends, which is a question specifically about your wallet rather than the capabilities of the model.

Here’s the reason for the curve.

The issue Fable 5 had in the marketplace was not that it was insufficient; it was that it burned through tokens finishing tasks, blowing past budget and generating unbudgeted charges. Fortune directly quoted enterprise customers complaining about the number of tokens the model used to complete a task. So the next iteration of Claude ships with a knob labelled spend less.

The numbers, and who produced them

On Frontier-Bench v0.1, Anthropic claims Opus 5 exceeds all other models on every metric and achieves more than double the score of Opus 4.8 while costing less per task. On CursorBench 3.2, it ties Fable 5’s peak performance using max effort per task at under half the cost. On ARC-AGI 3, which specifically tests its ability to solve novel problems, it scores three times higher than any other model on average. And on OSWorld 2.0, a computer use benchmark, they claim Opus 5 exceeds Fable 5’s best performance at under a third of the cost.

All numbers from Anthropic, using their harness, their runs, their footnotes.

I don’t believe they’re lying, but the last two years should have taught us that vendor benchmarks are marketing copy with data attached, and that we should wait for Artificial Analysis before trusting any of it. Which brings us to the customer benchmarks, which are harder to dismiss.

Harvey (legal), when asked to compare Opus 5’s performance to its predecessors using equivalent reasoning tokens, found it achieved equal quality at 26% fewer tokens than Opus 4.8 at max reasoning. A trading firm compared its performance on an internal benchmark against all prior opuses and found it to exceed all of them using only 1/7th the reasoning tokens and under half the latency. Zapier’s CEO compared the runs on AutomationBench and found Opus 5 to be the first model to exceed their leaderboard using the same token budget as any prior Claude, and to have achieved 100% on the sequence for preventing customer churn that no other model had managed to complete.

These are the numbers I’m choosing to believe in first, because they’re all essentially saying the same thing: you get the same answer for significantly fewer tokens.

There’s also an anecdote from the release that I keep returning to, which I think gets at the same point. On a Frontier-Bench task, the model was given a drawing of a machine part and asked to produce code to recreate the part in FreeCAD, but was deliberately denied any means of viewing the image. Opus 5 wrote its own computer vision pipeline to extract the geometry from the pixels and build the part. Anthropic notes that no other model solved the task in five attempts.

Another data point from the press release, presented as a story rather than a chart. A good story, also, because you might be looking for one if you’ve already read the fine print.

The classifiers are the product now.

This is the section that gets little attention, but it’s the one that fundamentally shifts the conversation for you. Anthropic chose not to train Opus 5 on cyber tasks, and found it substantially improved at cybersecurity as a side effect. It’s approaching Mythos 5 in its ability to find vulnerabilities, but significantly worse at exploiting them, which is the other side of the coin.

So the safeguards broadly fall on that line; Opus 5 will find bugs in your source code, but cannot scan binaries, perform penetration testing, or generate exploits. Safeguards that once fired in Claude’s API now fire with roughly 85% fewer tokens in Opus 5, according to Anthropic. When one of these classifiers does fire, the request is refused. In Claude, Claude Code, and Cowork, this means the request is sent to Opus 4.8 instead. In the API, automatic fallbacks are now a documented beta: requests that trigger them will see their traffic routed elsewhere, rather than being outright denied.

That, to me, represents a fundamental shift in how these companies operate and how you should think about them. If the request failed outright, if the model was shut down, it was a failure of the system. Now, that failure is rerouted and mitigated, which is not much different than the way the models are shifting towards mitigation rather than outright refusal when a guardrail is crossed.

This is the product of a company that evidently learned a particular lesson in June, and shipped improvements to this particular facet of Claude’s operation as infrastructure.

The same pattern follows the alignment updates; the automated behavioural audit scores Opus 5 at 2.3 average misaligned behaviour, the lowest of any recent model, and better than Opus 4.8, Sonnet 5, or Fable 5 in terms of general adherence to Claude’s Constitution. Internal metric, internal scale; nothing to report externally, but worth noting as a datapoint in and of itself.

What this actually means

Fourth iteration of Claude in under two years. Mythos 5, Fable 5, Sonnet 5, Opus 5. The competition has moved on, as they always do.

When a pair of models in the same class score within half a per cent of each other on a coding benchmark, the capability stops being the product; the bill and the failure modes become the product, and it becomes a question of the model checking its own work before telling you it’s done.

A maturing market, perhaps, rather than an arms race.

But the part I find most interesting is the part that connects back to June. Anthropic still recommends Fable 5 for the most demanding autonomous tasks, the things that run for days at a time. Mythos 5, the redacted version of Fable 5 without the safety controls, is better still at long-horizon offensive cyber tasks and biological research.

Mythos 5 is not available to the public, or even to most organisations. It’s distributed through Project Glasswing, the restricted partnership program I’ve written about previously.

So here, as of this week, is the state of the art for frontier intelligence: cheaper to run for longer stretches of time, and the capabilities that let it operate beyond the oversight of any external authority are gated through a separate process.

Anthropic will tell you this is responsible scaling, and they’re not wrong to think so. I’ve argued for it myself, several times.

It’s also a world where the most capable systems are distributed as allocation rather than purchase, and where those allocations are non-transparent. Both can be true, but only one appears in the press release.

Add SpaceRock as a preferred source on Google