DeepSeek
DeepSeek V4.1 Flash pricing and benchmarks
DeepSeek’s low-cost model with thinking and non-thinking modes, tool calls and vision input.
Input
$0.30
per 1M tokens
Cached input
$0.006
per 1M tokens
Output
$1.20
per 1M tokens
How capable is DeepSeek V4.1 Flash?
Epoch AI has not published capability scores for DeepSeek V4.1 Flash yet. See benchmarks for rated models.
What DeepSeek V4.1 Flash costs in practice
| Workload | Monthly cost |
|---|---|
Support chatbot 100,000 replies a month, ~1,500 input tokens (system prompt + history) and ~400 output tokens each. | $93.00 |
Document Q&A (RAG) 50,000 questions a month with ~8,000 tokens of retrieved context and ~600 output tokens each. | $156.00 |
Coding agent 10,000 agent steps a month, ~40,000 input tokens each (70% read from the prompt cache) and ~2,000 output tokens. | $61.68 |
Pricing details
- Prices shown are peak rates. Off-peak rates are half: peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday.
- Compatible with both OpenAI-format and Anthropic-format APIs.
Estimate your DeepSeek V4.1 Flash bill
- DeepSeek V4.1 FlashDeepSeek$156.00/mo$0.0031 / request
Standard-tier list prices. Long-context rates apply automatically where the provider publishes a threshold. Excludes cache-write surcharges, taxes, batch and volume discounts.