DeepSeek V4-Flash
A 284B-parameter MoE with only 13B active, which is why it runs fast and prices low. Notable for a 384K maximum output — far beyond the 128K most frontier models cap at — and for supporting both thinking and non-thinking modes, so you can switch reasoning off on easy requests rather than paying for it.
Pricing checked against provider documentation on . How we verify
deepseek-v4-flash reasoning open weights Input
$0.22
/ 1M tokens
Output
$0.66
/ 1M tokens
Context
1M
1,000,000 tokens
Max output
384K
tokens per request
Pricing caveat: Off-peak rate shown. Peak rates double ($0.44 / $1.32) during 01:00–04:00 and 06:00–10:00 UTC on weekdays. Superseded the earlier flat-rate pricing of $0.14 / $0.28 that many comparison sites still quote.
What it costs in practice
A workload of 10M input and 2M output tokens per month — roughly a moderately busy production assistant — runs $3.52/month on DeepSeek V4-Flash.
Where DeepSeek V4-Flash fits
- High-volume classification and extraction
- Very long generated outputs
- Budget coding agents
- Local deployment on modest hardware
Specifications
- API model ID
- deepseek-v4-flash
- Provider
- DeepSeek
- Model family
- DeepSeek V4
- Released
- June 2026
- Knowledge cutoff
- March 2026
- Context window
- 1,000,000 tokens
- Max output
- 384,000 tokens
- Native reasoning
- Yes
- Relative latency
- low
- Input modalities
- text
- Output modalities
- text
- Open weights
- MIT
- Cached input
- $0.007 / 1M tokens
- Batch discount
- —
Frequently asked
How much does DeepSeek V4-Flash cost?
DeepSeek V4-Flash costs $0.22 per million input tokens and $0.66 per million output tokens, with cached input at $0.01.
What is DeepSeek V4-Flash's context window?
DeepSeek V4-Flash accepts up to 1,000,000 tokens of context and can generate up to 384,000 output tokens per request.
Is DeepSeek V4-Flash a reasoning model?
Yes. DeepSeek V4-Flash performs native chain-of-thought before answering. Those thinking tokens are billed at the output rate, so budget above the sticker price.
What is DeepSeek V4-Flash best for?
The cheapest capable open-weight option. A 284B-parameter MoE with only 13B active, which is why it runs fast and prices low. Notable for a 384K maximum output — far beyond the 128K most frontier models cap at — and for supporting both thinking and non-thinking modes, so you can switch reasoning off on easy requests rather than paying for it.