DeepSeek Undercuts Astra on Output by 83x

At 10:58 pm on September 9, DeepSeek sent out a notice: V4.1 Flash was about to ship, and afterwards “all requests to the Pro model will be routed to V4.1 Flash and billed at Flash’s price”.
Less than 14 hours later, at 12:42 pm on September 10, a second email arrived — same subject, different decision. V4 Pro’s discontinuation was “postponed to 12:00 Beijing Time on September 14, 2026”, and the email carried a sentence the first one did not: “If you encounter any issues during your comparative testing between V4 Pro and V4.1 Flash, please do not hesitate to reach out to us with your feedback.”
A company pushing back the retirement of its own flagship less than fourteen hours after announcing it, and explicitly inviting users to report comparison problems, is not the sort of detail that makes it into a press release. Only the API customers who received both emails saw it. And September 14 at noon is this coming Monday.
The new prices: emails and official docs match line for line
The numbers first. Below is what DeepSeek’s official API documentation currently lists (USD per million tokens), which matches the tables in both emails exactly:
| Per 1M tokens | V4.1 Flash off-peak | V4.1 Flash peak | V4 Pro off-peak | V4 Pro peak |
|---|---|---|---|---|
| Input (cache hit) | $0.003 | $0.006 | $0.022 | $0.044 |
| Input (cache miss) | $0.15 | $0.30 | $0.66 | $1.32 |
| Output | $0.60 | $1.20 | $1.98 | $3.96 |
The new pricing took effect at 04:00 UTC on September 10, 2026. V4 Pro’s price is unchanged for as long as it remains in service — so until noon on September 14 you can still reach Pro, and still pay Pro rates.
Measured purely against DeepSeek’s own previous generation: output drops from $1.98 to $0.60 (3.3x), cache-miss input from $0.66 to $0.15 (4.4x), and cache hits from $0.022 to $0.003 (7.3x). The cache-hit column falls hardest, which matters most for agent-style workloads that resend the same prefix over and over.
Three things inside nine days
Laid out on a timeline, this price change lands in a very crowded stretch:
- September 1: Anthropic ships Claude Fable 5.1. Base rates unchanged at $10 input / $50 output, but cache reads drop from $1 per million tokens to $0.25.
- September 3-4: OpenAI launches GPT-6 Astra — limited preview on the 3rd, stable for paying users on the 4th.
- September 10: DeepSeek ships V4.1 Flash and announces that V4 Pro traffic will be billed at the new model’s price.
Nine days, three vendors. That is the only claim this section makes. Whether it amounts to “competitive pressure” is the subject of the last section, because that part is inference, not data.
Three-way price comparison
Public list prices for each vendor’s current flagship or workhorse model, in USD per million tokens:
| Model | Input | Output | Cache read |
|---|---|---|---|
| GPT-6 Astra | $10 | $50 | $1 |
| Claude Fable 5.1 | $10 | $50 | $0.25 |
| DeepSeek V4.1 Flash (off-peak) | $0.15 | $0.60 | $0.003 |
| DeepSeek V4.1 Flash (peak) | $0.30 | $1.20 | $0.006 |
The first thing worth noticing: GPT-6 Astra and Claude Fable 5.1 carry identical list prices, $10 and $50. Two companies, independent launches, entirely different architectures, landing on the same number — which tells you how thoroughly anchored pricing at this tier has become.
The second is the gap itself. On output, V4.1 Flash off-peak is 1/83 of those two. On input it is 1/67. On cache reads it is 1/83 of Fable 5.1, and against Astra’s $1 it is 1/333.
A concrete monthly workload makes it clearer. Assume 10 million input tokens per month (all cache misses) plus 2 million output tokens:
| Model | Monthly cost |
|---|---|
| GPT-6 Astra | $200.00 |
| Claude Fable 5.1 | $200.00 |
| DeepSeek V4 Pro (off-peak, before Sept 14) | $10.56 |
| DeepSeek V4.1 Flash (peak) | $5.40 |
| DeepSeek V4.1 Flash (off-peak) | $2.70 |
$200 against $2.70 — roughly 74x.
But price is not a like-for-like comparison
That table is striking, and easy to misread, so it is worth being precise about what it does not establish.
DeepSeek’s own wording is that V4.1 Flash “has comprehensively surpassed V4 Pro across all key metrics, including performance, cost, speed, and task completion time”. The comparison is against its own previous model — not Astra, not Fable 5.1. As of writing I could find no public third-party benchmark putting V4.1 Flash head to head with either.
So the correct reading is: DeepSeek executed a cheaper-and-better generational swap inside its own line-up, and its absolute prices sit far below the two American flagships. “74x cheaper” and “74x cheaper while being just as good” are different claims, and nothing here supports the second. Any decision to move production traffic should rest on measurements of your own workload, not on this table.
It is also not a same-tier comparison. Astra and Fable 5.1 are each company’s most capable model, while Flash is positioned — and named — as the lightweight option. The genuinely comparable matchup arrives when DeepSeek ships V4.1 Pro, which both emails describe as not yet released.
Peak hours are defined in UTC, and that matters
DeepSeek defines peak hours in UTC: 01:00–04:00 and 06:00–10:00, Monday to Friday, with everything else off-peak. If you work in UTC+8 — Beijing, Taipei, Singapore — that converts to:
09:00–12:00 and 14:00–18:00 local time, Monday to Friday.
In other words, peak pricing covers the entire working day, interrupted only by the lunch hour. Coding, running agents, and batch processing from the office all fall inside the double-rate window, while off-peak rates only arrive after hours, overnight, or at weekends.
Two practical consequences:
- Anything schedulable should not run during office hours. Data cleaning, batch summarisation, overnight regression suites — work that does not need an immediate answer — halves in price simply by moving past 18:00 or onto a weekend. In the worked example above, that is $5.40 becoming $2.70.
- Weekends are entirely off-peak. Peak hours are Monday to Friday only, so Saturday and Sunday bill at off-peak around the clock.
Worth noting: because Beijing time is also UTC+8, the “9:00-12:00, 14:00-18:00” written in DeepSeek’s Chinese documentation can be read directly as local time in that zone. The English emails state the window in UTC — read only that one and it is easy to be wrong by eight hours.
The bigger lever is cache hit rate, not scheduling
Rescheduling saves half. Do only that, and you leave a much larger sum on the table.
Look again at the table: V4.1 Flash charges $0.003 for a cache hit and $0.15 for a miss — a 50x spread, against just 2x between peak and off-peak. For the same 10 million input tokens per month, at off-peak rates:
| Cache hit rate | Monthly input cost |
|---|---|
| 0% | $1.500 |
| 50% | $0.765 |
| 80% | $0.324 |
| 90% | $0.177 |
| 95% | $0.104 |
Going from no caching to a 90% hit rate cuts input cost 8.5x; moving everything off-peak cuts it 2x. Do both by all means, but if time is short, fix caching first.
Cache hits depend on the prefix being byte-for-byte stable. The usual things that quietly break it: a datetime.now(), a per-request UUID, or a request id near the top of the system prompt; serialising tool definitions with unsorted JSON; adding or removing tools between requests. Any of those changes the prefix hash and re-prices everything after it. Put stable content first and volatile content last, then check that the cache-read token count the API returns is actually above zero — if it is zero across repeated requests, something is invalidating the cache silently.
What to do before noon on September 14
Only days remain before V4 Pro stops serving, and Beijing noon is noon across UTC+8. If your service currently points at deepseek-v4-pro:
- Confirm whether you are actually on Pro. If you call the model id
deepseek-v4-pro, requests after noon on September 14 will not fail — they will quietly run on V4.1 Flash instead. The model changes and your code raises nothing. A silent substitution is harder to notice than a 4xx. - Run your comparison while Pro still exists. This is exactly what the second email invites. Send your real prompts through both and judge the output yourself rather than taking the announcement’s word for it — especially if your prompts were tuned against Pro’s behaviour over months.
- Re-estimate cost rather than assuming savings. Unit prices fall (about 3.9x off-peak), but if the new model needs more turns or longer outputs to finish the same job, the realised saving is smaller than the headline. Measure cost per completed task, not per million tokens.
- Audit your caching. With a 50x spread between hit and miss, whether your prefix reliably caches matters far more to the bill than the 2x peak/off-peak difference.
How to read the “pressure” question
Which brings us to the tempting question: did Astra and Fable 5.1 force DeepSeek’s hand?
The honest answer is that neither email mentions a competitor even once, and neither does the documentation. The reading rests entirely on timing, and coincidence is not evidence. Model launches are scheduled months ahead; nine days apart may be nothing more than a collision.
What can be stated as fact is narrower: within a week of two American flagship launches, DeepSeek shipped a model it says beats its own previous generation on every metric; it cut prices substantially; it intended to redirect old-model traffic immediately, then reversed that in under fourteen hours to buy four more days, while publicly asking users to report problems from comparison testing.
That last point is the one I find most telling. The reversal says more than the price cut does. A decision made in full confidence rarely needs to be pushed back inside a day of the notice going out, with “let us know if you hit issues” appended. Whether that came from customer pushback, internal reservations, or simple courtesy is not something outsiders can know — but the sentence is theirs, not my inference.
A price war at the 1/83 level is good news for users. Just remember, before you move production traffic across, that the table compares prices, not capability.



