About DeepSeek V4 Pro
DeepSeek V4 Pro is the performance end of the V4 series — the same MIT-licensed, mixture-of-experts, compressed-attention design as its Flash sibling, scaled to a total parameter count in the trillion class with a correspondingly larger active set per token. DeepSeek positions it as the choice when reasoning depth matters more than response snappiness, with the same switchable thinking modes.
The weights are open, but at this scale that is a statement of principle more than a deployment plan; in practice Pro is consumed through DeepSeek's API, and that is how we treat it here.
Pro costs a few times Flash's rates while remaining far below Western frontier pricing — that gap is the entire commercial argument, and for many workloads it holds up. Both models move together on DeepSeek's peak and off-peak schedule: weekday business hours, stated in UTC on the pricing page, cost double the off-peak figure, and the callout above says which rate the numbers here reflect. The automatic cache-hit discount applies at a rate so far below standard input that prefix-heavy workloads see their input costs nearly vanish without any cache management. Budget primarily for output: thinking-mode deliberation arrives as output tokens, and hard problems can produce a lot of invisible reasoning. For pipelines that can wait, scheduling around the peak window is worth as much as any prompt optimisation.
Frequently asked questions
How much does the DeepSeek V4 Pro API cost?
DeepSeek V4 Pro costs $1.32 per million input tokens and $3.96 per million output tokens, with cached input at $0.044 per million (97% off). Note: Peak-hour rates (01:00–04:00 and 06:00–10:00 UTC, Mon–Fri); all rates are half price off-peak. Cached-input price is DeepSeek’s cache-hit rate.
How much does a typical chat message cost with DeepSeek V4 Pro?
A short chat message (about 500 input and 250 output tokens) costs roughly $0.0016 with DeepSeek V4 Pro. At 1,000 such requests a day, that is about $49.50 per month.
Is DeepSeek V4 Pro cheap compared to similar models?
DeepSeek V4 Pro has the #17 cheapest input rate of the 29 models we track. On a mixed workload (10K input / 2K output), it costs $0.021 per request, versus $0.015 for Gemini 3.8 Flash.
What does DeepSeek V4 Pro cost per 1,000 tokens?
Per 1,000 tokens, DeepSeek V4 Pro costs $0.0013 for input and $0.0040 for output, or $0.00004 for cached input. Providers quote rates per million tokens ($1.32/M in, $3.96/M out here), so divide by a thousand for the per-1K figure older pricing pages used.
How large a prompt can DeepSeek V4 Pro take?
DeepSeek V4 Pro accepts up to 977K tokens of input context and can generate up to 375K tokens of output per request. At its input rate, filling the entire 977K-token window costs about $1.32 per request before any output.
Related
Compare all models at once ·Count tokens in your prompt
All DeepSeek model costs: DeepSeek V4 Flash
Last updated 2026-09-03. Prices verified against DeepSeek's official pricing page; see the methodology.