
Anthropic’s cheapest route to frontier-level AI landed on 24 July, and on two of its own plans it landed without anyone choosing it. Claude Opus 5 is now the default model on Claude Max and the strongest model available on Claude Pro, priced at roughly £3.76 per million input tokens and £18.83 per million output tokens ($5 and $25, at about $1.33 to the pound). That is unchanged from its predecessor, Opus 4.8, and half the token price of Claude Fable 5, the company’s flagship.
The pricing is the story, because the capability gap it used to buy has almost closed. Independent testing published the same day by Artificial Analysis put Opus 5 at 61 on its Intelligence Index at maximum effort, a point ahead of Fable 5 on 60 and two ahead of OpenAI’s GPT-5.6 Sol on 59. Marketing teams that spent the past year building business cases around which vendor holds the top benchmark slot now face a narrower and more practical question: which setting of which model finishes a given job for the least money.
Claude Opus 5 costs the same as the model it replaces and half as much as the model it nearly matches. Fable 5 remains at roughly £7.53 and £37.65 per million input and output tokens ($10 and $50); Sonnet 5 sits well below both at about £1.51 and £7.53 ($2 and $10) on introductory pricing. Cache hits are discounted 90%, to around £0.38 ($0.50) per million tokens, and batch processing halves the bill again for work that does not need an immediate answer.
Anthropic offers the model in five effort settings — low, medium, high, xhigh and max — with a one-million-token context window and an optional Fast mode that runs about 2.5 times quicker for twice the base price. The effort setting is the lever that matters commercially, because it changes how many tokens the model spends before answering. Artificial Analysis found the range spans 407 Elo points on one knowledge-work benchmark, with output token use varying roughly eightfold between the lowest and highest settings.
On its own evaluations, Anthropic reports that Opus 5 more than doubles Opus 4.8’s score on the software-engineering benchmark Frontier-Bench v0.1 at a lower cost per task, lands within 0.5% of Fable 5’s peak CursorBench 3.2 score at half the cost, and scores three times the next-best model on ARC-AGI 3. The claim closest to marketing operations comes from Zapier, which says Opus 5 topped its AutomationBench leaderboard and ran a churn-prevention sequence end to end from a raw account-health workbook, flagging at-risk accounts and alerting owners where earlier models failed the task outright. Those are the company’s figures and a customer’s, generated on Anthropic’s harness. The independent numbers below are the ones worth planning against.
Artificial Analysis ranked Opus 5 first on AA-Briefcase, its benchmark for agentic knowledge work, which scores models on realistic tasks spanning thousands of input files and requiring research reports, presentations and spreadsheets as deliverables. At max effort the model scored 1,720 Elo against Fable 5’s 1,574, at a cost of about £13.39 per task ($17.79) versus £16.79 ($22.30) — a fifth cheaper for a better result. At high effort it still beat Fable 5, by 32 Elo, for £7.84 ($10.41), less than half the price.
The breakdown is more useful than the ranking. Opus 5’s gains came from rubric pass rate and analytical quality, where it scored 2,016 Elo at max effort, nearly 300 clear of Fable 5. On presentation quality it reached 1,628, roughly 40 Elo behind GPT-5.6 Sol. For a team producing finished decks and client-ready reports rather than analysis alone, that split is the number to weigh, and it points the other way to the headline.
Speed is the second qualification. Opus 5’s top three effort settings each averaged more than 25 minutes per AA-Briefcase task — 36.2, 34.3 and 25.7 minutes — against 24.1 minutes for Opus 4.8 at max. The model took 103 turns per task at max effort where Opus 4.8 took 55. A cheaper token is not a cheaper hour, and a workflow with a person waiting at the end of it costs more than the invoice shows.
Artificial Analysis disclosed that it supported Anthropic in evaluating the model ahead of release, and ran its Intelligence Index tests with Opus 4.8 fallback enabled. That is pre-release access on the vendor’s timetable rather than a cold audit, which is the ordinary condition of launch-day benchmarking and worth registering before the figures reach a business case. The Helm reported a similar caveat when OpenAI published benchmarks alongside GPT-5.6 on 9 July.
Claude Opus 5 moved backwards on factual reliability, and Anthropic’s announcement does not mention it. On AA-Omniscience, Artificial Analysis found the model improved 7 points on accuracy over Opus 4.8 but answered more often when it was uncertain, pushing its hallucination rate up 14 points to 50%. Its factual knowledge still trails Fable 5. For teams using the model for competitive research, market sizing or anything with a citation in it, that is the single most consequential figure in the release, and it argues for keeping the verification step that AI-assisted research is often bought to remove.
Anthropic’s own testing points at a different quality. Its automated behavioural audit scored Opus 5 at 2.3 for overall misaligned behaviour, the lowest of its recent models, and the company describes it as its most aligned model to date, adhering to Claude’s Constitution more closely than Opus 4.8, Sonnet 5 or Fable 5. Alignment and factual accuracy are not the same property. A model can be more careful about what it should do and still be more willing to guess.
The model a marketing team is actually talking to is also, increasingly, not the one it selected. Requests that trip Opus 5’s cybersecurity classifiers fall back to Opus 4.8 by default in Claude.ai, Claude Code and Claude Cowork, and biology-related requests blocked on Fable 5 now route to Opus 5 instead. Anthropic says Opus 5’s cyber classifiers should intervene around 85% less often than Fable 5’s, and that the model remains behind Mythos 5 on both biology research and offensive cybersecurity. The pattern is familiar: Fable 5 returned from its government-ordered suspension on 1 July with a filter that rerouted flagged requests to an older model. Teams auditing AI output for accuracy are auditing a routing decision as much as a model choice.
Anthropic’s introductory pricing on Sonnet 5, the cheaper model most teams use for routine drafting, runs until 31 August, after which it rises from about £1.51 and £7.53 per million input and output tokens to £2.26 and £11.29 ($2 and $10 to $3 and $15).