Polarison

DeepSeek

DeepSeek V4.1 Flash pricing and benchmarks

DeepSeek’s low-cost model with thinking and non-thinking modes, tool calls and vision input.

Input
$0.30
per 1M tokens
Cached input
$0.006
per 1M tokens
Output
$1.20
per 1M tokens

How capable is DeepSeek V4.1 Flash?

Epoch AI has not published capability scores for DeepSeek V4.1 Flash yet. See benchmarks for rated models.

What DeepSeek V4.1 Flash costs in practice

WorkloadMonthly cost
Support chatbot
100,000 replies a month, ~1,500 input tokens (system prompt + history) and ~400 output tokens each.
$93.00
Document Q&A (RAG)
50,000 questions a month with ~8,000 tokens of retrieved context and ~600 output tokens each.
$156.00
Coding agent
10,000 agent steps a month, ~40,000 input tokens each (70% read from the prompt cache) and ~2,000 output tokens.
$61.68

Pricing details

  • Prices shown are peak rates. Off-peak rates are half: peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday.
  • Compatible with both OpenAI-format and Anthropic-format APIs.

Estimate your DeepSeek V4.1 Flash bill

Standard-tier list prices. Long-context rates apply automatically where the provider publishes a threshold. Excludes cache-write surcharges, taxes, batch and volume discounts.