DeepSeek Quietly Ships V4 Pro, Targets Fable 5

閱讀中文版 →

DeepSeek Quietly Ships V4 Pro, Targets Fable 5

According to Taiwan’s United Daily News, DeepSeek pushed out V4 Pro’s official release late at night on August 12 in what the outlet called a “surprise” move — no blog post, no announcement, the version number just quietly ticked over to 0813. Independent AI commentator Simon Willison confirmed as much: he discovered the new version through OpenRouter, since there was nothing on DeepSeek’s own site announcing it. And the comparison DeepSeek chose to publish alongside this release goes straight at Anthropic’s most expensive flagship model: Claude Fable 5.

Who Fable 5 actually is

According to other tech outlets’ reporting, Claude Fable 5 is Anthropic’s “Mythos-class” model, released June 9, 2026, positioned a tier above Claude Opus 4.8, with a 1-million-token context window and up to 128,000 tokens of output. It’s priced at $10 per million input tokens and $50 per million output tokens — among the most expensive of any mainstream model on the market today. Choosing to benchmark a new flagship directly against the industry’s acknowledged top performer, and priciest option, is a fairly pointed statement on its own.

Is the gap 5.3%, or 2.8%

According to DeepSeek’s own published comparison, across 10 agent benchmarks, Claude Fable 5 leads by an average of 5.3%, with DeepSeek actually coming out ahead on two of them. But strip out Humanity’s Last Exam — the one benchmark where the gap balloons to 10.6% — and the average margin across the rest narrows to just 2.8%. That said, not every gap is small: on SWE-bench Verified specifically, Claude Fable 5 scores 96% against DeepSeek V4 Pro’s 80.6%, a 15.4-point difference that’s considerably wider than the overall average.

The real gap is the price

Price is where the real separation shows up. DeepSeek V4 Pro charges $0.435 per million input tokens (just $0.003625 on a cache hit) and $0.87 per million output tokens; Claude Fable 5 charges $10 and $50 respectively. On blended rates, that works out to roughly $30 versus $0.65 — about a 46x gap. United Daily News cites a concrete measurement from benchmarking firm Artificial Analysis: running the same benchmark task costs $3.15 on Fable 5, versus just $0.03 on DeepSeek’s lighter V4 Flash — a 99% reduction.

None of these numbers have actually been independently verified yet

Worth flagging plainly: every comparison figure above comes from DeepSeek’s own published results — no independent third party has run a full retest of the 0813 release yet. Willison also noticed something odder still: these performance numbers first surfaced in a WeChat group, got copied over to Reddit, were deleted there for being “low-effort content,” and eventually resurfaced as an ASCII-art table on Hacker News. That’s a strikingly informal path for benchmark data from a company of DeepSeek’s size, and a notable gap from the kind of communication you’d expect from a company with this much market weight behind it.

Willison also flagged another quirk he says he hasn’t seen from any other model: running the same prompt through V4 Pro at “low,” “medium,” and “high” reasoning settings produces noticeably different image outputs — something he called “a kind of difference I’ve not noticed from any other model.” Separately, while April’s preview build and July’s V4 Flash release both had their weights published openly on Hugging Face under an MIT license, whether the 0813 build will get the same open-weight treatment is something even Willison hasn’t been able to confirm yet.

Yuan pricing and other context

United Daily News also compiled DeepSeek’s yuan-denominated pricing for the Chinese market: ¥0.025 per million tokens for cached input, ¥3 for non-cached input, and ¥6 for output — compared to the previously released V4 Flash’s ¥0.02, ¥1, and ¥2 respectively. V4 Pro is priced above Flash across the board, consistent with its position as the more capable flagship model.

This rollout follows the same pattern DeepSeek has used for past releases: no ad campaign, no launch event, a new model quietly swapped in overnight — paired, this time, with a scorecard that directly challenges the most expensive model in the industry. However quiet the release itself was, the challenge it’s throwing down is anything but.

About the author

I’m Ryan, and I run RyanOps. My day job is software development and automation; here I track what changes in AI models, developer tools and software engineering, and write up hands-on notes from problems I have debugged and built myself.

About this site and the editorial process →