DeepSeek shipped DeepSeek-V4-Pro as its flagship this week and, at the same time, rewrote its price sheet. Caixin reports increases ranging from 50% to more than 1,100% depending on the model, the token type and the time of day. The new pricing takes effect at 16:00 UTC on Sunday, 16 August, and covers both V4-Flash and V4-Pro.
The mechanics are new for DeepSeek. Pricing is now split by peak and off-peak windows — peak running 9am to noon and 2pm to 6pm Beijing time. During peak hours, V4-Pro costs 0.3 yuan per million cached input tokens, 9 yuan per million uncached input tokens and 27 yuan per million output tokens. The old flat rates were 0.025, 3 and 6 yuan respectively. Off-peak is set at half the peak rate.
That cached-input line is the eye-catcher: 0.025 to 0.3 yuan is a twelvefold increase, which is where the 1,100% headline comes from. In dollar terms the peak cached rate is still about four cents per million tokens. DeepSeek remains cheap in absolute terms. It is no longer cheap relative to itself.
The direction of travel just inverted
This lands in the same week that US labs went the other way. OpenAI has cut pricing on GPT-5.6 Luna, Anthropic has positioned Claude Opus 5 at roughly half the price of its higher-end Fable 5, and Google shipped Gemini 3.7 Flash at half the previous Flash cost. The Financial Times reports that prices customers actually pay for leading US models have fallen materially since mid-July.
So the market now has premium Chinese models getting more expensive while American ones get cheaper — the exact opposite of the narrative that has driven AI policy debate for eighteen months.
Our take: Peak/off-peak pricing is not a pricing strategy, it is a capacity strategy. You charge more at 10am Beijing time because you cannot serve everyone at 10am Beijing time. DeepSeek is telling the market it is compute-constrained and has decided to monetise the constraint rather than eat it. That is a rational move for a company that can’t buy its way out with Nvidia hardware — and it hands US labs a talking point they did not have to pay for. If your cost model assumed Chinese inference prices only ever fall, rebuild it.
What this means if you’re buying tokens
Three practical consequences. First, time-of-day now matters: batch jobs that can wait until off-peak are worth half as much to run. Second, cache design gets more expensive to get wrong — the cached-input rate went up more, in percentage terms, than anything else on the sheet. Third, single-vendor inference is a risk position; a supplier that can reprice twelvefold with 48 hours’ notice is not a fixed cost.
What to watch
- Sunday, 16:00 UTC. Whether the increase lands as published, or gets softened after the reaction.
- Whether Moonshot, Z.ai and Alibaba follow. If they hold prices, DeepSeek loses share fast. If they follow, the cheap-Chinese-inference era is over as a category.
- V4-Flash. The cheap tier survives. How much traffic migrates down is the real demand test.
- DeepSeek’s own silicon. The company is building an in-house inference chip. Price hikes buy time for that to arrive.
