Polarison

Google

Gemini 3.5 Flash pricing and benchmarks

Google’s Gemini 3.5 Flash model, with thinking, tool use and a 1M-token context window.

Input
$1.50
per 1M tokens
Cached input
$0.15
per 1M tokens
Output
$9.00
per 1M tokens
Capability index
154.7
#12 of 16 tracked

How capable is Gemini 3.5 Flash?

Gemini 3.5 Flash scores 154.7 on the Epoch Capabilities Index (confidence interval 152.8157.0), ranking #12 of the 16 rated models we track.

BenchmarkTestsScore
GPQA DiamondScience reasoning
90.4%
FrontierMath (Tiers 1–3)Advanced math
62.8%
SimpleQA VerifiedFactual accuracy
66.2%
ARC-AGI-2Abstract reasoning
72.1%
DeepSWECoding
37.4%
APEX-AgentsAgentic work

Source: Epoch AI (as “Gemini 3.5 Flash”), best recorded result, CC BY 4.0. A dash means no published score. Compare all models

What Gemini 3.5 Flash costs in practice

WorkloadMonthly cost
Support chatbot
100,000 replies a month, ~1,500 input tokens (system prompt + history) and ~400 output tokens each.
$585.00
Document Q&A (RAG)
50,000 questions a month with ~8,000 tokens of retrieved context and ~600 output tokens each.
$870.00
Coding agent
10,000 agent steps a month, ~40,000 input tokens each (70% read from the prompt cache) and ~2,000 output tokens.
$402.00

Pricing details

  • Output price includes thinking tokens.
  • Context cache storage costs $1.00 per 1M tokens per hour.
  • Priority inference costs 1.8x the standard rate.

Estimate your Gemini 3.5 Flash bill

Standard-tier list prices. Long-context rates apply automatically where the provider publishes a threshold. Excludes cache-write surcharges, taxes, batch and volume discounts.