DeepSeek V4 Pro Ships, Price Hike Already Flagged

閱讀中文版 →

DeepSeek V4 Pro Ships, Price Hike Already Flagged

On August 12, 2026, DeepSeek’s flagship model V4 Pro wrapped up a preview period that lasted nearly four months, showing up as the official “V4-Pro-0813” release on OpenRouter’s model page and in DeepSeek’s own API documentation. But right as the model finally shipped, the company’s own notice already flagged what’s coming next: a “significant increase” in price.

Back to the April preview

According to DeepSeek’s own documentation, the V4 series first showed up as a preview on April 24, 2026, launching two models at once: V4-Pro, with 1.6 trillion total parameters and 49 billion active per inference, and the lighter V4-Flash, with 284 billion total parameters and 13 billion active. Both were released under an MIT open-source license, defaulted to a 1-million-token context window, and shipped with both “Thinking” and “Non-Thinking” modes, integrating into tool chains like Claude Code, OpenClaw, and OpenCode. Architecturally, DeepSeek paired token-wise compression with its own DeepSeek Sparse Attention (DSA) mechanism, and pre-trained the models on more than 32 trillion tokens.

What happened in between

According to other outlets tracking DeepSeek’s changelog during this period, on July 24, DeepSeek formally retired the legacy API aliases deepseek-chat and deepseek-reasoner — a deprecation schedule it had announced three months earlier. Then on July 31, things took an interesting turn: the smaller V4-Flash graduated first, entering public beta as “V4-Flash-0731.” DeepSeek’s own changelog described it as having “significantly enhanced agent capabilities, far exceeding” the still-in-preview V4-Pro; independently verified scores put this build at 82.7 on Terminal-Bench 2.1 and 54.4 on DeepSWE. As for the bigger V4-Pro, the company’s statement at the time was simply that the official release “will follow soon,” with no specific timeline given.

Pro finally shipped — here’s where it actually improved

It wasn’t until August 12 that V4-Pro-0813 closed that gap. Compared to the prior V3.2 generation, DeepSeek says that at maximum context settings, the new version cuts single-token inference compute down to just 27% and KV cache usage down to just 10%. On performance, the version DeepSeek calls V4-Pro-Max scored 80.6% on SWE-bench Verified, 93.5% pass@1 on LiveCodeBench, 90.1% pass@1 on GPQA Diamond, 87.5% on MMLU-Pro, and a Codeforces rating of 3,206.

Where it wins, where it doesn’t

Set against the competitive field, V4 Pro still trails GPT-5.4 on Terminal Bench 2.0 (67.9% vs. 75.1%). On SWE-bench Verified, its 80.6% trails Gemini 3.1 Pro and Claude Opus 4.6 by a hair — both of those score 80.8%, a gap of just 0.2 percentage points, essentially neck and neck. On LiveCodeBench, it comes out on top among the models compared. Overall, this is a model that’s nipping at the heels of the top competitors on coding-focused benchmarks, but hasn’t fully overtaken GPT-5.4 across the board.

The pricing looks great — for now

V4 Pro’s current API pricing carries over from the preview period: $0.435 per million input tokens on a cache miss, dropping to $0.003625 per million on a cache hit, and $0.87 per million output tokens — still cheap by the standards of comparable models. But DeepSeek has already stated plainly in its own notice that a “significant increase” is coming, and separately announced a peak-hour pricing policy (2x rates from 9am-12pm and 2pm-6pm Beijing time) that hasn’t taken effect yet. In other words, the price you’re seeing right now is likely the cheapest this model will ever be.

Open weights remain an open question

The weights from April’s preview release are still up on Hugging Face under the same MIT license. But DeepSeek hasn’t yet said whether this August GA release, the 0813 build, will get its own separate open-weight release — and for a company that’s built much of its identity around being the open-source alternative, whether that actually happens is probably worth watching just as closely as the benchmark scores.

About the author

I’m Ryan, and I run RyanOps. My day job is software development and automation; here I track what changes in AI models, developer tools and software engineering, and write up hands-on notes from problems I have debugged and built myself.

About this site and the editorial process →