DeepSeek V4 Flash API Cost Calculator

DeepSeek's DeepSeek V4 Flash: $0.44/M input · $1.32/M output · $0.014/M cached input. Price your own workload below.

Peak-hour rates (01:00–04:00 and 06:00–10:00 UTC, Mon–Fri); all rates are half price off-peak. Cached-input price is DeepSeek’s cache-hit rate.

Presets:

What DeepSeek V4 Flash costs in practice

WorkloadInputOutputCost
Short chat message500250$0.0006
10-page document + summary7,000500$0.0037
Agent / coding session50,00010,000$0.035
1M tokens in + 1M out1,000,0001,000,000$1.76

DeepSeek V4 Flash vs nearest rivals

Closest-priced alternatives on a mixed 10K-input / 2K-output request:

ModelInput $/MTokOutput $/MTokMixed request
DeepSeek V4 Flash$0.44$1.32$0.0070
GPT-5 mini$0.25$2$0.0065
Gemini 2.5 Flash$0.3$2.5$0.0080
Mistral Large 3$0.5$1.5$0.0080

About DeepSeek V4 Flash

DeepSeek V4 Flash is the efficiency-focused member of DeepSeek's V4 series: an MIT-licensed open-weight mixture-of-experts model whose active parameters per token are a small slice of its large total, paired with heavily compressed attention that keeps its million-token context practical. DeepSeek pitches it at fast, intuitive responses for routine work, with switchable thinking modes when a task needs actual deliberation.

It leads a double life by design — the same weights serve DeepSeek's low-cost API and anyone's own infrastructure — which is why it appears in both our cost and VRAM guides.

DeepSeek bills by the clock as well as by the token: its pricing page defines weekday peak hours in UTC terms — matching its own business day — during which every rate is double the off-peak figure, and the callout above states which of the two this page uses. Even at peak the rates sit well below hosted models of comparable scale from Western providers, and the caching story is the better one: cache hits are priced automatically at a small fraction of the input rate with no cache management or storage fees, so repetitive prefixes get the discount without engineering for it. Thinking modes bill through the same token accounting, so deliberate answers cost more via length rather than a separate rate; shifting batch work into off-peak hours is the other lever.

Frequently asked questions

How much does the DeepSeek V4 Flash API cost?

DeepSeek V4 Flash costs $0.44 per million input tokens and $1.32 per million output tokens, with cached input at $0.014 per million (97% off). Note: Peak-hour rates (01:00–04:00 and 06:00–10:00 UTC, Mon–Fri); all rates are half price off-peak. Cached-input price is DeepSeek’s cache-hit rate.

How much does a typical chat message cost with DeepSeek V4 Flash?

A short chat message (about 500 input and 250 output tokens) costs roughly $0.0006 with DeepSeek V4 Flash. At 1,000 such requests a day, that is about $16.50 per month.

Is DeepSeek V4 Flash cheap compared to similar models?

DeepSeek V4 Flash has the #8 cheapest input rate of the 29 models we track. On a mixed workload (10K input / 2K output), it costs $0.0070 per request, versus $0.0065 for GPT-5 mini and $0.0080 for Gemini 2.5 Flash.

What does DeepSeek V4 Flash cost per 1,000 tokens?

Per 1,000 tokens, DeepSeek V4 Flash costs $0.00044 for input and $0.0013 for output, or $0.00001 for cached input. Providers quote rates per million tokens ($0.44/M in, $1.32/M out here), so divide by a thousand for the per-1K figure older pricing pages used.

How large a prompt can DeepSeek V4 Flash take?

DeepSeek V4 Flash accepts up to 1024K tokens of input context and can generate up to 375K tokens of output per request. At its input rate, filling the entire 1024K-token window costs about $0.461 per request before any output.

Related

Compare all models at once ·Count tokens in your prompt · Run DeepSeek V4 Flash locally — VRAM requirements · DeepSeek V4 Flash vs GPT-4o mini head-to-head

All DeepSeek model costs: DeepSeek V4 Pro

Last updated 2026-09-03. Prices verified against DeepSeek's official pricing page; see the methodology.