Gemini 3.5 Flash pricing and benchmarks
Google’s Gemini 3.5 Flash model, with thinking, tool use and a 1M-token context window.
Input
$1.50
per 1M tokens
Cached input
$0.15
per 1M tokens
Output
$9.00
per 1M tokens
Capability index
154.7
#12 of 16 tracked
How capable is Gemini 3.5 Flash?
Gemini 3.5 Flash scores 154.7 on the Epoch Capabilities Index (confidence interval 152.8–157.0), ranking #12 of the 16 rated models we track.
| Benchmark | Tests | Score |
|---|---|---|
| GPQA Diamond | Science reasoning | 90.4% |
| FrontierMath (Tiers 1–3) | Advanced math | 62.8% |
| SimpleQA Verified | Factual accuracy | 66.2% |
| ARC-AGI-2 | Abstract reasoning | 72.1% |
| DeepSWE | Coding | 37.4% |
| APEX-Agents | Agentic work | — |
Source: Epoch AI (as “Gemini 3.5 Flash”), best recorded result, CC BY 4.0. A dash means no published score. Compare all models
What Gemini 3.5 Flash costs in practice
| Workload | Monthly cost |
|---|---|
Support chatbot 100,000 replies a month, ~1,500 input tokens (system prompt + history) and ~400 output tokens each. | $585.00 |
Document Q&A (RAG) 50,000 questions a month with ~8,000 tokens of retrieved context and ~600 output tokens each. | $870.00 |
Coding agent 10,000 agent steps a month, ~40,000 input tokens each (70% read from the prompt cache) and ~2,000 output tokens. | $402.00 |
Pricing details
- Output price includes thinking tokens.
- Context cache storage costs $1.00 per 1M tokens per hour.
- Priority inference costs 1.8x the standard rate.
Estimate your Gemini 3.5 Flash bill
- Gemini 3.5 FlashGoogle$870.00/mo$0.02 / request
Standard-tier list prices. Long-context rates apply automatically where the provider publishes a threshold. Excludes cache-write surcharges, taxes, batch and volume discounts.