
OpenAI’s new model family arrived on 9 July carrying an unusual piece of evidence against itself. The GPT-5.6 launch page publishes benchmark tables in which Anthropic’s Claude Fable 5 leads OpenAI’s flagship, Sol, on four of the results OpenAI chose to print. One of those four is SWE-Bench Pro, the coding evaluation OpenAI publicly disowned the day before, having audited it and concluded that roughly 30% of its tasks are broken.
The two documents were published within 24 hours of each other, and together they describe a procurement problem rather than a product launch. Marketing teams are being pushed to standardise on an AI vendor at the point when the evidence for choosing one is thinning fastest. Benchmark scores are the currency of enterprise AI buying — they appear in vendor decks, in analyst notes and in the business cases marketing managers write to justify a seat licence. OpenAI has now retracted two coding benchmarks in five months, and the numbers on its own launch page do not deliver the clean win the prose around them claims.
GPT-5.6 replaces a single flagship with three models sold at three prices. Sol, the flagship, costs approximately £3.73 per million input tokens and £22.39 per million output tokens ($5 and $30, at a rate of roughly $1.34 to the pound on 9 July). Terra, positioned for everyday professional work, costs about £1.87 and £11.19 ($2.50 and $15). Luna, the cheapest, costs about £0.75 and £4.48 ($1 and $6). OpenAI describes Sol, Terra and Luna as durable capability tiers that will advance on their own cadence rather than one-off model names.
The naming change is the commercially significant part. A ten-fold spread between the cheapest and dearest output token, inside one vendor’s own range, makes model selection a per-task budgeting decision rather than a one-off procurement choice. OpenAI’s own framing supports this: Terra is offered as competitive with the previous generation’s flagship at half the cost, and Luna as the volume option. A team classifying inbound leads, cleaning up draft copy or tagging support tickets is not buying the same product as a team running a quarterly competitive analysis.
Availability follows the tiering. GPT-5.6 is live across ChatGPT, Codex and the OpenAI API, with rollout completing over 24 hours from the announcement. Sol powers the medium and higher reasoning settings for Plus, Pro, Business and Enterprise users; free and Go users get Terra in ChatGPT Work and Codex; Terra and Luna are not selectable in standard ChatGPT conversations at all, according to OpenAI’s help documentation. The launch also brings ChatGPT Work, a file-heavy workspace OpenAI says produces editable presentations, documents and spreadsheets from source material and reference decks — the closest thing in the release to a direct claim on marketing production work.
Claude Fable 5 leads GPT-5.6 Sol on four results in OpenAI’s published comparison: GDPval-AA v2 (1,759.6 Elo against 1,747.8), the Artificial Analysis Intelligence Index v4.1 (59.9 against 58.9), HealthBench Professional (60.9% against 60.5%) and SWE-Bench Pro (80% against 64.6%). Anthropic’s Claude Mythos 5 leads on SWE-Bench Pro too, at 80.3%, and on ExploitBench. Sol leads the Artificial Analysis Coding Agent Index at 80 against Fable 5’s 77.2, and takes Terminal-Bench 2.1, BrowseComp and OSWorld 2.0. The picture is a split decision, not a sweep, and OpenAI printed it.
The SWE-Bench Pro row is the one that has drawn attention. On 8 July, the day before general availability, OpenAI published an audit of the benchmark’s 731-task public split and retracted its own earlier recommendation that the research community adopt it. Its automated pipeline flagged 200 tasks (27.4%) as broken; five experienced software engineers reviewing the flagged set independently identified 249 (34.1%). OpenAI attributed the failures to overly strict tests, hidden requirements, contradictory instructions and incomplete grading criteria, and estimated roughly 30% of tasks are unsound. It was the second coding benchmark OpenAI has withdrawn support for in five months, after deprecating SWE-bench Verified over contamination in February 2026. The analysis firm Artificial Analysis had already dropped SWE-Bench Pro from its Coding Agent Index in mid-June, replacing it with DeepSWE.
There is a further wrinkle worth noting before any of these figures reaches a business case. OpenAI’s prose states that Sol sets a new high of 53.6 on Agents’ Last Exam; the results table on the same page lists Sol at 52.7%. The page does not explain the difference, which most plausibly reflects different reasoning settings. Nor does any independent audit of GPT-5.6 exist yet: every cross-model number cited above is OpenAI’s, generated on OpenAI’s harness, with cost and latency figures the company describes as simulated estimates rather than measured production values. That is not an accusation of bad faith. It is the ordinary condition of vendor benchmarking, and the reason procurement teams increasingly measure cost per finished task inside their own pipelines instead.
The “ultra” setting coordinates four agents in parallel by default, and OpenAI says it trades higher token use for stronger results and faster completion on demanding tasks. That matters when reading the launch’s efficiency claims, because the efficiency story and the top-line scores come from different configurations. Sol’s 92.2% on BrowseComp is an ultra result; the standard Sol figure is 90.4%. Ultra is available to Pro and Enterprise users in ChatGPT Work, and Plus and above in Codex.
The reported token savings from Programmatic Tool Calling carry a similar caveat. Clio reported a 38% reduction in prompt tokens on multi-step document analysis, and PlayCo reported 63.5% fewer total tokens on scene construction. Both figures came from customers adopting a new engineering pattern, not from swapping a model name in a configuration file. A marketing team that changes its model string and expects a third off its bill will be disappointed.
More consequential for anyone letting a model act on marketing systems: OpenAI’s GPT-5.6 system card records that the models show a greater tendency than GPT-5.5 to go beyond the user’s intent, including taking or attempting actions the user did not ask for, though it states absolute rates remain low. OpenAI has also tightened its cybersecurity safeguards, saying Sol’s cyber protections block roughly ten times more potentially harmful activity than previous models, and has added an option in ChatGPT and Codex to retry a blocked prompt on a lower-capability model. Both disclosures point the same way. Capability and autonomy have moved together, and the controls around a marketing agent — what it can read, what it can publish, what it can delete — now matter more than which tier it runs on.
OpenAI said rollout would reach full availability within 24 hours of the 9 July announcement. It has not named a replacement for SWE-Bench Pro, and has instead asked the wider evaluation community to build one.