AI & Technology

Cheaper AI models cost more in nearly a third of tests, researchers find

Written by
Full Name
October 6, 2026
A study of eight frontier models by Stanford, Berkeley, Carnegie Mellon and Microsoft researchers found listed token prices often fail to predict the real bill, as a separate survey shows just 11% of organisations forecast AI spend to within 10%.

A lower price per token can mean a higher bill. Researchers at Stanford University, the University of California, Berkeley, Carnegie Mellon University and Microsoft Research ran eight frontier reasoning models through 6,877 tasks and found that in 32% of head-to-head comparisons, the model with the lower listed price cost more to complete the work.

The finding, reported by The Wall Street Journal, lands as companies struggle to budget for AI at all. A survey of 396 organisations by Mavvrik and Benchmarkit found only 11% could forecast their AI spending to within 10%, down from 15% a year earlier. For marketing teams now paying for agents, content generation and research tools on usage-based plans, the two numbers describe the same problem: a rate card prices tokens, while a marketing team pays for finished tasks.

What did the price reversal study find?

The study, “The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More”, was posted to arXiv in March 2026 and revised in May. Its authors, who include Lingjiao Chen, Ion Stoica, Matei Zaharia and James Zou, compared eight models across 12 benchmarks covering competition maths, science questions, coding and multi-step agent tasks. The line-up included OpenAI’s GPT-5.4 and GPT-5.4 Mini, Google’s Gemini 3.1 Pro and Gemini 3 Flash, and Anthropic’s Claude Opus 4.7 and Claude Haiku 4.5.

Of 336 pairwise cost comparisons, 106 showed a reversal, and the worst gap reached 28 times. The clearest case involved Gemini 3 Flash. Its listed price is 80% below GPT-5.4’s, yet across all tasks it cost 38% more to run: roughly £520 against £380 ($705 against $509).

The cause lies in how reasoning models work. They generate “thinking” tokens before they answer, and those tokens are billed. A cheaper model that reasons for longer, or takes more turns inside an agent loop, can spend its per-token discount and more. On a single question, the paper found one model could use 900% more thinking tokens than another. Chen told the Journal that price alone should not be used to infer which model is actually cheaper.

Why do identical AI tasks cost different amounts each time?

Cost per task is not fixed even when the model and the prompt stay the same. The researchers found that running an identical query repeatedly produced thinking-token counts that varied by up to 9.7 times between the cheapest and most expensive run. The paper treats the actual cost of a single prompt on a single model as a random variable.

The Journal illustrated the swing with an example from the researchers’ data. Given the same prompt, Gemini 3.1 Pro finished in 85 steps for about 75p ($1). Gemini 3 Flash, the lighter and cheaper model, took nearly 1,000 steps, spent about £10 ($14), and still failed. The Journal acknowledged the case could be an outlier but said such results also occur in ordinary use.

Google, in a statement to the Journal, said some fluctuation at the level of individual prompts is inherent to AI and averages out across large, varied workloads. It pointed to spending caps and flexible pricing as controls customers can use. Google, Anthropic and OpenAI have all released newer models since the testing, so the specific pairings may not hold for current versions. The mechanism behind them, variable reasoning effort billed per token, is unchanged.

How are organisations and marketing teams handling AI costs?

Mavvrik’s survey, run with Benchmarkit in April and May 2026, suggests most organisations track AI spend without controlling it. Of respondents, 98% said they track AI costs and 95% have formal AI budgets. Yet 62% reported cost surprises that forced material changes to the business, and 25% delayed or cancelled AI initiatives because of cost. Agents are a particular blind spot: 98% run agentic workloads, but only 36% include them in cost reporting. Mavvrik sells AI cost-governance software, so the findings come from a company with a commercial interest in the problem. The sample spanned technology, financial services, retail and manufacturing.

Marketing has met the issue first through its agencies. Digiday reported in March that agencies were split on how to bill for tokens. Merge passes metered costs to clients case by case, Big Spaceship treats compute as a production line item, and RPA absorbs the cost while work remains experimental. One Coca-Cola campaign cited in the report needed 70,000 prompts and millions of tokens.

For in-house teams, the research points to practical steps. Model choice becomes a cost decision that can be tested: running a representative batch of a team’s own tasks through two or three models gives a truer price than any rate card. Spending caps, per-task cost monitoring and clear guidance on which model suits which job, the remedy the Journal suggests, give finance a firmer number to plan against. The paper’s authors go further and call on vendors to provide transparent per-request cost reporting.

The researchers list predicting the cost of a single AI request as an open research problem, one that model providers have yet to solve.

Subscribe to our newsletter

By subscribing you agree to with our Privacy Policy
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Share article

Recommended Reading