DeepSeek V4 Flash API pricing
DeepSeek's fast, ultra-cheap tier with a 1M-token context window — half price off-peak.
Input
$0.44
per 1M tokens
Cached input
$0.014
per 1M tokens
Output
$1.32
per 1M tokens
Context window
1M
max output 384K
Prices last synced Aug 20, 2026 from the LiteLLM community price catalog — the same catalog Marginal prices real usage events against, refreshed daily. Hand-checked against the provider's official pricing page on Aug 20, 2026. Spot an error? Tell us.
Pricing notes
- Rates shown are peak-hour (01:00–04:00 and 06:00–10:00 UTC). ALL other hours are off-peak at exactly half price: $0.22 input / $0.007 cached / $0.66 output per 1M.
- Input is billed by cache outcome — a cache hit costs ~1/30th of a miss. Thinking and non-thinking modes cost the same.
- 1M-token context, up to 384K output tokens.
Also sold by
Each host prices the model independently — same weights, different bill.
| Host | Input / 1M | Output / 1M | Context |
|---|---|---|---|
| DeepSeek (first-party) | $0.44 | $1.32 | 1M |
| Fireworks AI | $0.14 | $0.28 | 1M |
What would DeepSeek V4 Flash cost you?
Enter your workload — add other models to compare the same traffic across providers.
Your monthly workload
Cache hit rate = share of input tokens served from the provider's prompt cache, billed at the model's cached-input price. Estimates use standard per-token rates — batch discounts, service tiers, and long-context surcharges are not applied.
| Model | Input | Cached input | Output | Per request | Monthly |
|---|---|---|---|---|---|
| DeepSeek V4 Flash | $264 | $5.60 | $330 | $0.00120 | $600 |
Compare DeepSeek V4 Flash
More from DeepSeek
Tracking DeepSeek V4 Flash in production?
Marginal prices every call your app actually makes — cached tokens, dated snapshots, model switches included — and slices spend by customer or feature. One track() call to integrate.