DeepSeek Price Hike: 3 Days Left at the Old Rate

閱讀中文版 →

DeepSeek Price Hike: 3 Days Left at the Old Rate

A plainly-worded notice landed in my inbox at 16:55 on August 14, 2026: DeepSeek is changing its API billing. Cross-checking it against DeepSeek’s own pricing documentation turned up something bigger than a simple price increase — the company is introducing peak/off-peak dual-tier billing, and there’s a timing gap that works in developers’ favor: the new prices don’t take effect until 00:00 Beijing time on August 17, but the new model is already live.

You’re Already Running the New Model

Open DeepSeek’s API docs and you’ll see that deepseek-v4-pro currently maps to the underlying model version DeepSeek-V4-Pro-0813 — the version number alone tells you this build shipped on August 13. In other words, anyone calling deepseek-v4-pro right now is already getting the latest pre-hike model capability, while still being billed at the old rate. That “new model, old price” window is exactly what this article wants to flag: between now and 00:00 on August 17, there’s a last chance to run the same model at a noticeably lower cost.

The Full Old-vs-New Pricing Table

Here’s the complete old and new billing schedule as published in DeepSeek’s official docs (in RMB per million tokens):

Old pricing (current, valid until 8/17)

ModelInput (cache hit)Input (cache miss)OutputConcurrency limit
deepseek-v4-flash¥0.02¥1¥22500
deepseek-v4-pro¥0.025¥3¥6500

New pricing (effective 8/17 00:00, split by peak/off-peak)

ModelTime slotInput (cache hit)Input (cache miss)Output
deepseek-v4-flashOff-peak¥0.05¥1.5¥4.5
deepseek-v4-flashPeak¥0.10¥3.0¥9.0
deepseek-v4-proOff-peak¥0.15¥4.5¥13.5
deepseek-v4-proPeak¥0.30¥9.0¥27.0

DeepSeek’s official documentation defines peak hours as 9:00–12:00 and 14:00–18:00 Beijing time daily, with everything else counted as off-peak; off-peak pricing is fixed at exactly half of peak pricing. That means the same model, the same billing line item, will now carry two different rates within a single day — a clear shift away from the flat, single-rate billing DeepSeek users have been used to.

Breaking Down the Actual Multipliers

Lining up the old and new numbers line by line shows the increases are far from uniform — the steepest jump lands on cache-hit input, which used to be the cheapest line item by far:

  • deepseek-v4-pro / cache-hit input: ¥0.025 → off-peak ¥0.15 (6x) / peak ¥0.30 (12x), the single largest jump in the entire schedule.
  • deepseek-v4-flash / cache-hit input: ¥0.02 → off-peak ¥0.05 (2.5x) / peak ¥0.10 (5x).
  • deepseek-v4-pro / output: ¥6 → off-peak ¥13.5 (2.25x) / peak ¥27 (4.5x).
  • deepseek-v4-flash / output: ¥2 → off-peak ¥4.5 (2.25x) / peak ¥9 (4.5x).
  • deepseek-v4-pro / cache-miss input: ¥3 → off-peak ¥4.5 (1.5x) / peak ¥9 (3x).
  • deepseek-v4-flash / cache-miss input: ¥1 → off-peak ¥1.5 (1.5x) / peak ¥3 (3x).

It’s worth noting that cache-hit input pricing was already extremely cheap to begin with (¥0.025 per million tokens), so while the multiplier looks alarming, its real impact on a total bill is much smaller than output and cache-miss input — the two line items that typically make up the bulk of any API bill. Looking at output pricing, the figure most developers will actually feel, the peak-hour increase comes to 4.5x — that’s the number that matters for most real-world usage. Factor in the edge case of cache-hit input and the headline “largest increase” climbs to 12x.

Why Introduce Peak/Off-Peak Pricing At All?

DeepSeek’s documentation doesn’t spell out the business reasoning behind the change, but the billing structure itself is a reasonable clue: peak/off-peak dual-tier pricing is a long-standing practice in cloud computing and utilities, designed to shift demand away from peak hours and even out server load to raise overall resource utilization. For DeepSeek, rather than absorbing every spike in traffic during peak hours by scaling up expensive GPU capacity, using price signals to nudge batch jobs and non-time-sensitive requests toward off-peak hours is the cheaper lever to pull. That also means workflows without strict latency requirements — batch data processing, offline content generation — can still land costs close to today’s rates simply by rescheduling to off-peak hours going forward.

What Developers Should Do Right Now

The window closes at 00:00 Beijing time on August 17, and teams already using or evaluating the DeepSeek API have two things worth checking in the next few days. First, audit which billing line items you actually rely on most heavily — if your usage leans hard on deepseek-v4-pro cache hits or output tokens, the post-hike cost jump will be the most noticeable, so budget for it now. Second, if your workload can tolerate latency, this is a good moment to evaluate shifting scheduled jobs to off-peak hours (18:00 to 9:00 daily, plus 12:00–14:00) — the savings compound meaningfully over time. Whether it’s worth front-loading extra usage between now and 8/17 comes down to your account’s concurrency limits (2500 for flash, 500 for pro) — however short the window is, it’s still the only chance to get the new model at the old price.

About the author

I’m Ryan, and I run RyanOps. My day job is software development and automation; here I track what changes in AI models, developer tools and software engineering, and write up hands-on notes from problems I have debugged and built myself.

About this site and the editorial process →