AI & Technology

DeepSeek raises V4 API prices from 16 August and adds peak-hour billing

Written by
Full Name
August 15, 2026
DeepSeek will charge almost five times more for V4-Flash output at peak hours from 16:00 UTC on 16 August, and split its rate card into peak and off-peak windows, turning the hour a job runs into a cost decision.

The vendor that set the AI market’s price floor is about to raise it. DeepSeek, the Chinese lab whose rock-bottom rates the rest of the industry has been measured against, said on 13 August that it will raise prices across its V4 family from 16:00 UTC on Sunday 16 August and introduce separate peak and off-peak rates.

The change lands on marketing teams that have spent the past year shifting repetitive work onto cheap models: bulk product descriptions, first-draft ad variants, transcript summarisation, lead enrichment, sentiment classification. Those workloads were costed against an assumption that token prices only ever fall, even as large enterprises capped staff AI use when token bills outran budgets. DeepSeek’s new rate card is the clearest signal yet that the floor exists, and that it moves.

What changes on 16 August?

DeepSeek’s new rate card roughly triples input prices and more than quadruples output prices at peak hours. V4-Flash output rises from £0.21 ($0.28) per million tokens to £0.98 ($1.32) at peak and £0.49 ($0.66) off-peak. Cache-miss input on the same model goes from £0.10 ($0.14) to £0.33 ($0.44) at peak and £0.16 ($0.22) off-peak. The larger V4-Pro moves from £0.64 ($0.87) per million output tokens to £2.93 ($3.96) at peak and £1.47 ($1.98) off-peak. Conversions use the mid-market rate of £1 to $1.35 on 13 August; DeepSeek bills in dollars, so a sterling budget carries the currency movement as well as the rise.

The company published the change alongside the general availability of DeepSeek-V4-Pro, which also brought adjustable reasoning-effort settings across both models and native support for OpenAI’s Responses API. DeepSeek said it was revising prices “to allocate resources more reasonably”, and presented the split rate card as a way to encourage more flexible workload scheduling.

The rise was signalled a week earlier. On 6 August the company told developers on its platform that a significant increase was coming and advised them to plan accordingly, without naming figures or a date. Bloomberg reported that the increase comes as DeepSeek prepares for a potential initial public offering.

Why does the hour a job runs now change the bill?

DeepSeek’s peak windows run from 01:00 to 04:00 and 06:00 to 10:00 UTC, with every other hour billed off-peak at half the peak rate. In British Summer Time that puts peak at 02:00 to 05:00 and 07:00 to 11:00, which places the UK working morning inside a peak window and the rest of the working day outside one.

The practical effect is that a batch job timed to finish before the team logs on now costs twice what the same job costs after 11:00. For interactive work, where a marketer is prompting a tool directly, the hour is not controllable and the peak rate applies. For queued work it is entirely controllable. Overnight enrichment runs, weekly content-generation batches, monthly citation-monitoring sweeps and scheduled summarisation pipelines can all move into off-peak hours without anyone noticing the difference in the output.

Caching cuts the bill further, and the structure of it is unchanged. DeepSeek charges a fraction of the standard input rate for cache hits, which matters most for workloads that resend the same system prompt, brand guidelines or reference document on every call. A content pipeline that sends a 4,000-token style guide with each of 2,000 requests pays close to full price for that context once rather than 2,000 times.

Teams buying AI through a martech vendor rather than direct API access will not see this rate card at all. What they may see over the following quarters is the vendor’s own pricing move, or a margin quietly absorbing the change.

Does DeepSeek still undercut OpenAI and Anthropic?

DeepSeek remains cheaper than its main Western rivals after the rise, by a narrower margin at peak. V4-Flash output at £0.98 ($1.32) still sits roughly 95% below Anthropic’s Claude Opus 5, listed at $25 per million output tokens, and about 97% below Claude Fable 5 at $50. Moonshot’s Kimi K3 lists at $15.

Analysts reading the new card describe the advantage as intact but conditional. Sanchit Vir Gogia, chief analyst at Greyhound Research, said that at peak, against the right comparator, DeepSeek’s price advantage disappears and in places inverts, but that “the schedule’s own clock and cache hand most of it back to any buyer paying attention”. Off-peak, he said, V4-Flash is marginally more expensive on input and 45% cheaper on output than OpenAI’s GPT-5.6 Luna.

Mark Tauschek, vice-president of research fellowships at Info-Tech Research Group, said the increase removes V4-Flash’s price advantage over Luna at peak rates but not off-peak, after OpenAI cut Luna’s API pricing by 80% in late July. V4-Pro, he said, keeps its advantage over OpenAI’s mid-tier reasoning model Terra even at peak rates.

Headline token rates are not the whole cost. Artificial Analysis, the San Francisco benchmarking firm, measured V4-Flash at roughly three cents per test run in early August, against 86 cents for Kimi K3, $1.86 for GPT-5.6 Sol and $3.15 for Claude Fable 5. That measure counts the tokens a model actually consumes to finish a task, which is the figure that reaches an invoice.

The new rates take effect at 16:00 UTC on 16 August. DeepSeek has not published rates beyond that card, or said whether the peak windows will be reviewed once demand settles.

Subscribe to our newsletter

By subscribing you agree to with our Privacy Policy
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Share article

Recommended Reading