Polarison

Google

Gemini 3.8 Flash pricing and benchmarks

Google’s newest Flash model, with thinking, tool use and a 1M-token context window.

Input
$0.75
per 1M tokens
Cached input
$0.075
per 1M tokens
Output
$3.75
per 1M tokens
Capability index
156.5
#6 of 16 tracked

How capable is Gemini 3.8 Flash?

Gemini 3.8 Flash scores 156.5 on the Epoch Capabilities Index (confidence interval 154.6162.0), ranking #6 of the 16 rated models we track. It is a best-value pick: no cheaper model we track scores higher.

BenchmarkTestsScore
GPQA DiamondScience reasoning
93.9%
FrontierMath (Tiers 1–3)Advanced math
68.4%
SimpleQA VerifiedFactual accuracy
69.7%
ARC-AGI-2Abstract reasoning
DeepSWECoding
73.8%
APEX-AgentsAgentic work

Source: Epoch AI (as “Gemini 3.8 Flash”), best recorded result, CC BY 4.0. A dash means no published score. Compare all models

What Gemini 3.8 Flash costs in practice

WorkloadMonthly cost
Support chatbot
100,000 replies a month, ~1,500 input tokens (system prompt + history) and ~400 output tokens each.
$262.50
Document Q&A (RAG)
50,000 questions a month with ~8,000 tokens of retrieved context and ~600 output tokens each.
$412.50
Coding agent
10,000 agent steps a month, ~40,000 input tokens each (70% read from the prompt cache) and ~2,000 output tokens.
$186.00

Pricing details

  • Introductory price through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, $7.50 output per 1M tokens.
  • Output price includes thinking tokens.
  • Priority inference costs 1.8x the standard rate.

Estimate your Gemini 3.8 Flash bill

Standard-tier list prices. Long-context rates apply automatically where the provider publishes a threshold. Excludes cache-write surcharges, taxes, batch and volume discounts.