On 6 August 2026 DeepSeek warned developers that API prices would rise in the near term and that the increase would be significant. It gave no figures and no date. We logged the question on the watchlist in the plainest form available: does DeepSeek raise its API pricing? At the time the company was charging USD 0.14 per million input tokens, which made it the cheapest capable model anyone could call.
It did. The change was published in the API changelog on 13 August 2026 and took effect at 16:00 UTC on 16 August 2026, alongside the general release of the V4 model family. The watchlist resolves YES.
The new structure is not simply a higher number. DeepSeek has abandoned flat pricing entirely in favour of peak and off-peak billing, with peak hours running 01:00 to 04:00 and 06:00 to 10:00 UTC and off-peak rates set at exactly half the peak rates. The current published table gives V4-Flash cache-miss input at USD 0.44 per million at peak and USD 0.22 off-peak, with output at USD 1.32 and USD 0.66. V4-Pro sits at USD 1.32 and USD 0.66 for cache-miss input and USD 3.96 and USD 1.98 for output. Cache hits remain very cheap, at USD 0.007 to USD 0.044 per million depending on model and hour.
Measured against the old flat rates, the increases run from roughly 50 percent to more than elevenfold, depending entirely on which model, which token type and which hour of the day. Output tokens took the heaviest rise. The company’s stated reason is resource allocation — the tiered structure is designed to push flexible workloads out of the two congested windows.
What it means is that the floor under the market has moved. DeepSeek’s flat USD 0.14 input rate was the number every other vendor’s pricing was implicitly compared against, and for eighteen months it functioned as a ceiling on what anyone else could charge for commodity inference. That anchor is gone. It is worth noting the sequence: DeepSeek raised prices on 16 August, six days after Anthropic cancelled a scheduled rise on Sonnet 5. The two moves point in opposite directions and it would be a mistake to read either as the trend.
For anyone running batch or overnight workloads, the off-peak rate is the number to plan against, and the eight peak hours are narrow enough that a scheduler can avoid them entirely. For interactive products serving European or Asian business hours, avoidance is not available, and the effective rise is closer to the top of the range than the bottom. Cache discipline matters far more than it did — the gap between a cache hit and a cache miss on V4-Pro is now a factor of thirty.
We are leaving a successor question open: does any vendor undercut the pre-August DeepSeek floor of USD 0.14 per million input tokens before 31 December 2026? Settlement is a published rate on a vendor’s own pricing page for a model with a context window of at least 128,000 tokens.