DeepSeek V4 ships in two open-weight sizes β€” Pro (1.6 trillion parameters, 49 billion active) and Flash (284 billion parameters, 13 billion active) β€” both MIT-licensed with a 1M-token context window. Since its April 2026 launch, DeepSeek has changed its pricing at least four times, most notably introducing peak-hour surge pricing on August 16, 2026 that quadruples output-token rates during two daily windows. Because those windows are set in UTC to match a Beijing working day almost exactly, they create a real asymmetry: US business hours land almost entirely in the cheaper off-peak window, while Asia-Pacific’s own daytime absorbs the full surge.

The actual pricing timeline, dated β€” because no single snapshot tells the story

DateWhat changed
April 24, 2026V4-Pro and V4-Flash launch, alongside a 75% promotional discount on V4-Pro
May 22, 2026DeepSeek makes the 75% V4-Pro discount permanent β€” no expiry date announced
June 30, 2026A leaked internal notice reveals plans for peak-hour surge pricing tied to the full V4 release
July 24, 2026Legacy deepseek-chat and deepseek-reasoner endpoints (from the V3/R1 generation) retire
August 13, 2026V4-Pro reaches general availability; Bloomberg reports the surge-pricing plan is confirmed
August 16, 2026Peak-hour pricing goes live β€” V4-Pro and V4-Flash output rates quadruple during peak windows
September 7, 2026Independent trackers confirm the new off-peak rates reflect the post-surge structure

Six real pricing events in under five months on the same model family is the concrete version of the caution our DeepSeek V4 vs GLM-5.3 comparison and our reasoning-token pricing guide both raise more generally: “permanent” and “no expiry date” language on any AI vendor’s pricing page has a shorter shelf life than it sounds. Check DeepSeek’s own live pricing page before budgeting, not this table or any other snapshot.

What surge pricing actually costs, in real numbers

ModelPre-surge rate (output)Peak rate (output)Off-peak rate (output)
V4-Pro~$0.87 per million$3.96 per million$1.98 per million
V4-Flash~$0.28 per million$1.32 per million$0.66 per million

Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday; every other hour is off-peak at exactly half the peak rate, per DeepSeek’s own current documentation. Even at the surcharged peak rate, DeepSeek remains far cheaper than Western flagship models β€” this is a company adding a floor and a ceiling to prices that were only ever falling, not a company becoming expensive by the standards of the broader market, as coverage of the change puts it plainly.

The timezone quirk nobody’s spelled out plainly

Convert DeepSeek’s UTC peak windows to Beijing time and they land at 9:00 AM–12:00 PM and 2:00 PM–6:00 PM β€” a standard Beijing working day, almost to the minute. The practical effect: a startup calling the API from San Francisco during its own normal 9-to-5 lands almost entirely in DeepSeek’s discounted off-peak window, while a startup in Shanghai working its own normal hours pays the full peak surcharge on the same model, from the same company. At a moment when US and EU policymakers are actively debating export restrictions and procurement bans aimed at Chinese open-weight models over competitive and security concerns, DeepSeek’s own rate card currently charges its own regional daytime users more than it charges customers on the other side of the debate. DeepSeek states its own reasoning is “better distribution of resources” and steadier service rather than profit β€” a claim worth noting as the company’s own framing, not an independently audited justification.

The lever that matters more than the clock

Before rearranging your workflow around DeepSeek’s peak windows, know that the surcharge applies to cache-hit input tokens too, but the absolute number involved is tiny β€” roughly half a cent per million tokens on V4-Pro even at peak, according to one detailed cost breakdown. Aggressive prompt caching still cuts your bill far more than timing requests around the surge windows ever will. If your workload is genuinely latency-tolerant β€” batch document processing, embeddings refreshes, non-urgent evaluation runs β€” scheduling it to avoid the 06:00–10:00 UTC window removes most of the surcharge without any other change to how you use the API.

What the model actually offers beyond the pricing story

  • Two sizes with a clear division of labor. V4-Pro is positioned for complex reasoning, software engineering, and multi-agent orchestration; V4-Flash is the faster, lower-latency variant, often deployed as a sub-agent handling simpler commands within DeepSeek’s own agent framework rather than as a standalone flagship.
  • A tunable reasoning-effort parameter. DeepSeek exposes thinking_effort with three settings β€” low, high, and max β€” controlling how much compute the model spends reasoning before it answers, directly trading off latency and cost the same way our reasoning-token pricing guide covers for other providers’ equivalent settings.
  • DeepSeek Harness, an open-source agent framework. Described by DeepSeek as modular β€” “everything is a plugin” β€” and natively optimized for both V4 models, it’s a genuine alternative for teams building agent workflows who want to avoid locking into a closed vendor’s own harness.
  • Genuine hardware flexibility. V4 has been adapted to run on Huawei’s Ascend chips at performance parity with Nvidia GPU deployments, a real option for anyone concerned about Nvidia supply constraints or export exposure, and a point of real divergence from GLM-5.3, whose smaller Flash variant still requires 4x H200 GPUs to run.

Treat the launch benchmarks the same way you’d treat any vendor’s own numbers

DeepSeek’s cited scores for V4-Pro-Max β€” 93.5 on LiveCodeBench Pass@1, a 3206 Codeforces rating β€” come from the company’s own initial release coverage, not an independent leaderboard, and the tracker reporting them explicitly flags that vendor-run scaffolds routinely score higher than independent harnesses reproduce. This is the same pattern that showed up in Meta’s Muse Spark 1.3 launch and Alibaba’s Qwen benchmarking this year β€” a vendor’s own scorecard is a starting point, not a verdict. Our comparison against GLM-5.3 covers this in more depth alongside GLM’s own unresolved benchmark disputes, the same caution our coding-agent harness guide applies across the board.

Who should actually use it

Cost-sensitive teams running batchable, latency-tolerant workloads get the most out of DeepSeek V4 β€” schedule the heavy lifting outside the UTC peak windows and the pricing becomes genuinely difficult to beat. Teams needing an unrestricted, genuinely open license, or wanting to avoid Nvidia-dependent infrastructure, have real reasons to choose it over closed alternatives. Teams building latency-sensitive, customer-facing products serving Asia-Pacific business hours specifically should model the peak surcharge into their unit economics rather than assuming the headline off-peak rate applies to their actual traffic pattern. And regardless of your use case, weigh this against our broader guide to free vs. paid AI tools and our guide to cutting AI costs for small teams rather than picking a model on price alone.

What is DeepSeek V4 and what does it cost?

It’s an open-weight, MIT-licensed model family from DeepSeek, shipping in Pro and Flash sizes with a 1M-token context window. Pricing depends on the time of day: V4-Pro runs $1.98 per million output tokens off-peak and $3.96 at peak; V4-Flash runs $0.66 off-peak and $1.32 at peak.

Why does DeepSeek’s pricing keep changing?

DeepSeek made a 75% discount permanent in May 2026, then introduced a new peak-hour surge structure in August that quadrupled output rates during specific daily windows. Both were real, dated changes reported accurately at the time β€” check the company’s live pricing page rather than trusting any comparison snapshot, including this one.

What are DeepSeek’s peak and off-peak hours, and why those specific windows?

Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, which converts almost exactly to a standard Beijing working day. Off-peak covers every other hour at half the peak rate.

Does DeepSeek’s surge pricing affect US-based businesses more or less than Asia-based ones?

Less, in practice. Because the peak windows are set to Beijing business hours, US working hours fall almost entirely into DeepSeek’s cheaper off-peak window, while Asia-Pacific daytime usage absorbs the full surcharge.

Can I reduce my DeepSeek bill without worrying about timing?

Yes β€” prompt caching matters far more than the clock. The peak surcharge applies to cached tokens too, but the absolute cost difference is negligible; aggressive caching cuts your bill more than scheduling requests around peak hours does.

Is DeepSeek V4 open source, and can I self-host it?

Yes, both V4-Pro and V4-Flash are released under the MIT license with weights available on Hugging Face, and the model has been adapted to run on Huawei’s Ascend chips at parity with Nvidia GPUs. Self-hosting at meaningful scale is still a serious infrastructure commitment, particularly for the larger Pro variant.

Shurah is the founder of AI Tools Daily, tracking pricing, licensing and policy changes across AI tools so readers can make decisions without wading through marketing claims themselves.