# Polarison: AI API pricing and capability reference

> Polarison (https://polarison.com) compares official AI API prices and Epoch AI capability scores. Prices verified September 14, 2026; benchmarks retrieved September 15, 2026. Cite as: Polarison, “AI API pricing and capability dataset”, https://polarison.com/data/models.json. Benchmark data: Epoch AI, “Capabilities & Benchmarking”, epoch.ai/benchmarks (CC BY 4.0).

## Key facts
- Prices last verified against official provider pages: September 14, 2026.
- Lowest blended price (3 input : 1 output): GPT-5.6 Luna (OpenAI), $0.45 per 1M tokens.
- Most capable tracked model by Epoch Capabilities Index: GPT-6 Astra (OpenAI), ECI 166.3.
- Highest DeepSWE coding score: GPT-6 Astra, 74.1%.
- Best-value models (no cheaper tracked model scores higher on ECI): GPT-5.6 Luna, Gemini 3.8 Flash, GPT-5.6 Terra, GPT-5.6 Sol, Claude Opus 5, GPT-6 Astra.

## Prices (USD per 1M tokens, standard tier)
Sorted by blended price = (3 × input + output) ÷ 4.

| Model | Provider | API model ID | Input | Cached input | Output | Blended | Context window | Max output |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| GPT-5.6 Luna | OpenAI | gpt-5.6-luna | $0.20 | $0.02 | $1.20 | $0.45 | 1.05M tokens | 128K tokens |
| DeepSeek V4.1 Flash | DeepSeek | deepseek-flash | $0.30 | $0.006 | $1.20 | $0.525 | 1M tokens | 384K tokens |
| Gemini 3.5 Flash-Lite | Google | gemini-3.5-flash-lite | $0.30 | $0.03 | $2.50 | $0.85 | 1.05M tokens | 66K tokens |
| Gemini 2.5 Flash | Google | gemini-2.5-flash | $0.30 | $0.03 | $2.50 | $0.85 | 1.05M tokens | 66K tokens |
| Gemini 3.8 Flash | Google | gemini-3.8-flash | $0.75 | $0.075 | $3.75 | $1.50 | 1.05M tokens | 66K tokens |
| Grok 4.3 | xAI | grok-4.3 | $1.25 | $0.20 | $2.50 | $1.563 | 1M tokens | — |
| DeepSeek V4 Pro | DeepSeek | deepseek-v4-pro | $1.32 | $0.044 | $3.96 | $1.98 | 1M tokens | 384K tokens |
| Claude Haiku 4.5 | Anthropic | claude-haiku-4-5-20251001 | $1.00 | $0.10 | $5.00 | $2.00 | 200K tokens | 64K tokens |
| Grok 4.6 | xAI | grok-4.6 | $2.00 | $0.50 | $6.00 | $3.00 | 500K tokens | — |
| Gemini 3.5 Flash | Google | gemini-3.5-flash | $1.50 | $0.15 | $9.00 | $3.375 | 1.05M tokens | 66K tokens |
| Gemini 2.5 Pro | Google | gemini-2.5-pro | $1.25 | $0.125 | $10.00 | $3.438 | 1.05M tokens | 66K tokens |
| Claude Sonnet 5 | Anthropic | claude-sonnet-5 | $2.00 | $0.20 | $10.00 | $4.00 | 1M tokens | 128K tokens |
| GPT-5.6 Terra | OpenAI | gpt-5.6-terra | $2.00 | $0.20 | $12.00 | $4.50 | 1.05M tokens | 128K tokens |
| Gemini 3.1 Pro | Google | gemini-3.1-pro-preview | $2.00 | $0.20 | $12.00 | $4.50 | 1.05M tokens | 66K tokens |
| GPT-5.6 Sol | OpenAI | gpt-5.6-sol | $4.00 | $0.40 | $20.00 | $8.00 | 1.05M tokens | 128K tokens |
| Claude Opus 5 | Anthropic | claude-opus-5 | $5.00 | $0.50 | $25.00 | $10.00 | 1M tokens | 128K tokens |
| GPT-6 Astra | OpenAI | gpt-6-astra | $10.00 | $1.00 | $50.00 | $20.00 | 1.05M tokens | 128K tokens |
| Claude Fable 5.1 | Anthropic | claude-fable-5-1 | $10.00 | $0.25 | $50.00 | $20.00 | 1M tokens | 128K tokens |

## Capability (Epoch AI, CC BY 4.0)
Epoch Capabilities Index (ECI) with confidence interval, plus benchmark scores in percent. “—” means no published score.

| Model | ECI | ECI range | GPQA Diamond | FrontierMath (Tiers 1–3) | SimpleQA Verified | ARC-AGI-2 | DeepSWE | APEX-Agents | Best value |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| GPT-6 Astra | 166.3 | 163.0–171.9 | 94.4% | 93.7% | 75.6% | 95.0% | 74.1% | 46.7% | yes |
| Claude Fable 5.1 | 164.5 | 161.4–168.3 | — | 90.2% | 70.8% | 90.0% | — | 47.4% | no |
| Claude Opus 5 | 162.3 | 159.9–165.7 | 91.8% | 85.6% | 59.9% | 90.4% | 73.6% | 43.5% | yes |
| GPT-5.6 Sol | 161.8 | 159.6–165.2 | 91.3% | 89.1% | 69.7% | 92.5% | 72.7% | 40.0% | yes |
| GPT-5.6 Terra | 159.1 | 157.0–161.8 | 91.1% | 86.0% | 43.2% | 83.9% | 69.6% | — | yes |
| Gemini 3.8 Flash | 156.5 | 154.6–162.0 | 93.9% | 68.4% | 69.7% | — | 73.8% | — | yes |
| Grok 4.6 | 156.3 | 154.8–158.9 | 92.0% | 66.0% | 49.3% | 67.1% | 67.5% | 41.2% | no |
| GPT-5.6 Luna | 156.3 | 154.1–158.9 | 88.8% | 82.1% | 41.0% | 59.5% | 67.2% | — | yes |
| Claude Sonnet 5 | 156.2 | 153.6–158.6 | 87.4% | 65.6% | 33.7% | — | 53.8% | 32.5% | no |
| DeepSeek V4 Pro | 155.5 | 153.8–157.8 | 88.9% | 64.6% | 52.9% | 61.3% | — | — | no |
| Gemini 3.1 Pro | 155.0 | 152.7–157.6 | 92.6% | 59.6% | 73.5% | 77.1% | 11.7% | 33.5% | no |
| Gemini 3.5 Flash | 154.7 | 152.8–157.0 | 90.4% | 62.8% | 66.2% | 72.1% | 37.4% | — | no |
| Gemini 2.5 Pro | 145.3 | 143.8–146.8 | 80.4% | 24.6% | — | 4.9% | — | 6.6% | no |
| Gemini 3.5 Flash-Lite | 145.1 | 142.8–146.8 | 77.8% | 26.0% | — | 10.3% | — | — | no |
| Claude Haiku 4.5 | 142.4 | 139.8–144.1 | 61.6% | — | 13.2% | 4.0% | — | 8.9% | no |
| Gemini 2.5 Flash | 140.5 | 138.5–141.9 | — | — | — | — | — | 1.8% | no |

Not yet rated by Epoch AI: Grok 4.3, DeepSeek V4.1 Flash.

## Monthly cost for typical workloads
- Support chatbot: 100,000 replies a month, ~1,500 input tokens (system prompt + history) and ~400 output tokens each.
- Document Q&A (RAG): 50,000 questions a month with ~8,000 tokens of retrieved context and ~600 output tokens each.
- Coding agent: 10,000 agent steps a month, ~40,000 input tokens each (70% read from the prompt cache) and ~2,000 output tokens.

| Model | Support chatbot | Document Q&A (RAG) | Coding agent |
| --- | --- | --- | --- |
| GPT-5.6 Luna | $78.00 | $116.00 | $53.60 |
| DeepSeek V4.1 Flash | $93.00 | $156.00 | $61.68 |
| Gemini 3.5 Flash-Lite | $145.00 | $195.00 | $94.40 |
| Gemini 2.5 Flash | $145.00 | $195.00 | $94.40 |
| Gemini 3.8 Flash | $262.50 | $412.50 | $186.00 |
| Grok 4.3 | $287.50 | $575.00 | $256.00 |
| DeepSeek V4 Pro | $356.40 | $646.80 | $249.92 |
| Claude Haiku 4.5 | $350.00 | $550.00 | $248.00 |
| Grok 4.6 | $540.00 | $980.00 | $500.00 |
| Gemini 3.5 Flash | $585.00 | $870.00 | $402.00 |
| Gemini 2.5 Pro | $587.50 | $800.00 | $385.00 |
| Claude Sonnet 5 | $700.00 | $1,100 | $496.00 |
| GPT-5.6 Terra | $780.00 | $1,160 | $536.00 |
| Gemini 3.1 Pro | $780.00 | $1,160 | $536.00 |
| GPT-5.6 Sol | $1,400 | $2,200 | $992.00 |
| Claude Opus 5 | $1,750 | $2,750 | $1,240 |
| GPT-6 Astra | $3,500 | $5,500 | $2,480 |
| Claude Fable 5.1 | $3,500 | $5,500 | $2,270 |

Estimates use standard list prices and exclude cache-write surcharges, taxes, batch and volume discounts.

## Model details

### GPT-6 Astra (OpenAI)
OpenAI’s most capable model, built for the hardest end-to-end reasoning and coding work.
- API model ID: gpt-6-astra
- Price: $10.00 input / $50.00 output per 1M tokens; cached input $1.00
- Long-context rates: $20.00 in / $75.00 out on long-context requests
- Batch discount: 50% off
- Context window: 1.05M tokens; max output: 128K tokens
- Epoch Capabilities Index: 166.3
- Cache writes are billed at $12.50 per 1M tokens.
- Fast mode costs 2x the standard rate; Flex costs half.
- Regional (data residency) endpoints add a 10% uplift.
- Official pricing: https://developers.openai.com/api/docs/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/gpt-6-astra

### GPT-5.6 Sol (OpenAI)
OpenAI’s flagship GPT-5.6 model for complex professional work.
- API model ID: gpt-5.6-sol
- Price: $4.00 input / $20.00 output per 1M tokens; cached input $0.40
- Long-context rates: $8.00 in / $30.00 out on long-context requests
- Batch discount: 50% off
- Context window: 1.05M tokens; max output: 128K tokens
- Epoch Capabilities Index: 161.8
- Promotional pricing, available at least through November 21, 2026.
- Cache writes are billed at $5 per 1M tokens.
- Regional (data residency) endpoints add a 10% uplift.
- Official pricing: https://developers.openai.com/api/docs/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/gpt-5-6-sol

### GPT-5.6 Terra (OpenAI)
The GPT-5.6 model that balances intelligence and cost.
- API model ID: gpt-5.6-terra
- Price: $2.00 input / $12.00 output per 1M tokens; cached input $0.20
- Long-context rates: $4.00 in / $18.00 out on long-context requests
- Batch discount: 50% off
- Context window: 1.05M tokens; max output: 128K tokens
- Epoch Capabilities Index: 159.1
- Cache writes are billed at $2.50 per 1M tokens.
- Fast mode costs 2x the standard rate; Flex costs half.
- Official pricing: https://developers.openai.com/api/docs/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/gpt-5-6-terra

### GPT-5.6 Luna (OpenAI)
The GPT-5.6 model optimized for cost-sensitive, high-volume workloads.
- API model ID: gpt-5.6-luna
- Price: $0.20 input / $1.20 output per 1M tokens; cached input $0.02
- Long-context rates: $0.40 in / $1.80 out on long-context requests
- Batch discount: 50% off
- Context window: 1.05M tokens; max output: 128K tokens
- Epoch Capabilities Index: 156.3
- Cache writes are billed at $0.25 per 1M tokens.
- Fast mode costs 2x the standard rate; Flex costs half.
- Official pricing: https://developers.openai.com/api/docs/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/gpt-5-6-luna

### Claude Fable 5.1 (Anthropic)
Anthropic’s model for demanding reasoning and long-horizon agentic work.
- API model ID: claude-fable-5-1
- Price: $10.00 input / $50.00 output per 1M tokens; cached input $0.25
- Long-context rates: —
- Batch discount: 50% off
- Context window: 1M tokens; max output: 128K tokens
- Epoch Capabilities Index: 164.5
- Cache hits cost 2.5% of the base input price, versus 10% on other Claude models.
- Cache writes cost $12.50 (5-minute cache) or $20 (1-hour cache) per 1M tokens.
- Uses Anthropic’s newer tokenizer, which produces roughly 30% more tokens for the same text.
- US-only inference (inference_geo: "us") costs 1.1x.
- Official pricing: https://platform.claude.com/docs/en/about-claude/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/claude-fable-5-1

### Claude Opus 5 (Anthropic)
Anthropic’s model for complex agentic coding and enterprise work.
- API model ID: claude-opus-5
- Price: $5.00 input / $25.00 output per 1M tokens; cached input $0.50
- Long-context rates: —
- Batch discount: 50% off
- Context window: 1M tokens; max output: 128K tokens
- Epoch Capabilities Index: 162.3
- Cache writes cost $6.25 (5-minute cache) or $10 (1-hour cache) per 1M tokens.
- Uses Anthropic’s newer tokenizer, which produces roughly 30% more tokens for the same text.
- US-only inference (inference_geo: "us") costs 1.1x.
- Official pricing: https://platform.claude.com/docs/en/about-claude/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/claude-opus-5

### Claude Sonnet 5 (Anthropic)
Anthropic’s best combination of speed and intelligence.
- API model ID: claude-sonnet-5
- Price: $2.00 input / $10.00 output per 1M tokens; cached input $0.20
- Long-context rates: —
- Batch discount: 50% off
- Context window: 1M tokens; max output: 128K tokens
- Epoch Capabilities Index: 156.2
- The $2 / $10 launch price is now the standard price; the planned rise to $3 / $15 will not happen.
- Cache writes cost $2.50 (5-minute cache) or $4 (1-hour cache) per 1M tokens.
- Uses Anthropic’s newer tokenizer, which produces roughly 30% more tokens for the same text.
- Official pricing: https://platform.claude.com/docs/en/about-claude/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/claude-sonnet-5

### Claude Haiku 4.5 (Anthropic)
Anthropic’s fastest model, with near-frontier intelligence.
- API model ID: claude-haiku-4-5-20251001
- Price: $1.00 input / $5.00 output per 1M tokens; cached input $0.10
- Long-context rates: —
- Batch discount: 50% off
- Context window: 200K tokens; max output: 64K tokens
- Epoch Capabilities Index: 142.4
- Cache writes cost $1.25 (5-minute cache) or $2 (1-hour cache) per 1M tokens.
- Uses Anthropic’s previous tokenizer.
- Official pricing: https://platform.claude.com/docs/en/about-claude/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/claude-haiku-4-5

### Gemini 3.8 Flash (Google)
Google’s newest Flash model, with thinking, tool use and a 1M-token context window.
- API model ID: gemini-3.8-flash
- Price: $0.75 input / $3.75 output per 1M tokens; cached input $0.075
- Long-context rates: —
- Batch discount: 50% off
- Context window: 1.05M tokens; max output: 66K tokens
- Epoch Capabilities Index: 156.5
- Introductory price through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, $7.50 output per 1M tokens.
- Output price includes thinking tokens.
- Priority inference costs 1.8x the standard rate.
- Official pricing: https://ai.google.dev/gemini-api/docs/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/gemini-3-8-flash

### Gemini 3.5 Flash (Google)
Google’s Gemini 3.5 Flash model, with thinking, tool use and a 1M-token context window.
- API model ID: gemini-3.5-flash
- Price: $1.50 input / $9.00 output per 1M tokens; cached input $0.15
- Long-context rates: —
- Batch discount: 50% off
- Context window: 1.05M tokens; max output: 66K tokens
- Epoch Capabilities Index: 154.7
- Output price includes thinking tokens.
- Context cache storage costs $1.00 per 1M tokens per hour.
- Priority inference costs 1.8x the standard rate.
- Official pricing: https://ai.google.dev/gemini-api/docs/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/gemini-3-5-flash

### Gemini 3.5 Flash-Lite (Google)
Google’s lowest-cost Gemini 3.5 model for high-volume work.
- API model ID: gemini-3.5-flash-lite
- Price: $0.30 input / $2.50 output per 1M tokens; cached input $0.03
- Long-context rates: —
- Batch discount: 50% off
- Context window: 1.05M tokens; max output: 66K tokens
- Epoch Capabilities Index: 145.1
- The same input price applies to text, image, video and audio.
- Output price includes thinking tokens.
- Official pricing: https://ai.google.dev/gemini-api/docs/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/gemini-3-5-flash-lite

### Gemini 3.1 Pro (Google)
Google’s Gemini 3.1 Pro model, currently offered as a preview.
- API model ID: gemini-3.1-pro-preview (preview)
- Price: $2.00 input / $12.00 output per 1M tokens; cached input $0.20
- Long-context rates: $4.00 in / $18.00 out above 200K prompt tokens
- Batch discount: —
- Context window: 1.05M tokens; max output: 66K tokens
- Epoch Capabilities Index: 155.0
- Prompts over 200K tokens are billed at $4 input / $18 output per 1M tokens.
- No free tier on the Gemini API.
- Output price includes thinking tokens.
- Official pricing: https://ai.google.dev/gemini-api/docs/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/gemini-3-1-pro

### Gemini 2.5 Pro (Google)
Google’s previous-generation Pro model.
- API model ID: gemini-2.5-pro
- Price: $1.25 input / $10.00 output per 1M tokens; cached input $0.125
- Long-context rates: $2.50 in / $15.00 out above 200K prompt tokens
- Batch discount: —
- Context window: 1.05M tokens; max output: 66K tokens
- Epoch Capabilities Index: 145.3
- Prompts over 200K tokens are billed at $2.50 input / $15 output per 1M tokens.
- Output price includes thinking tokens.
- Official pricing: https://ai.google.dev/gemini-api/docs/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/gemini-2-5-pro

### Gemini 2.5 Flash (Google)
Google’s previous-generation Flash model.
- API model ID: gemini-2.5-flash
- Price: $0.30 input / $2.50 output per 1M tokens; cached input $0.03
- Long-context rates: —
- Batch discount: —
- Context window: 1.05M tokens; max output: 66K tokens
- Epoch Capabilities Index: 140.5
- Audio input costs $1.00 per 1M tokens; text, image and video cost $0.30.
- Output price includes thinking tokens.
- Official pricing: https://ai.google.dev/gemini-api/docs/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/gemini-2-5-flash

### Grok 4.6 (xAI)
xAI’s flagship model for code and everything else, with agentic tool calling and configurable reasoning.
- API model ID: grok-4.6
- Price: $2.00 input / $6.00 output per 1M tokens; cached input $0.50
- Long-context rates: $4.00 in / $12.00 out above 200K prompt tokens
- Batch discount: —
- Context window: 500K tokens; max output: —
- Epoch Capabilities Index: 156.3
- Requests with 200K or more prompt tokens are billed at $4 input / $1 cached / $12 output for all tokens.
- Web Search and X Search tools cost $5 per 1,000 calls.
- Official pricing: https://docs.x.ai/docs/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/grok-4-6

### Grok 4.3 (xAI)
An earlier Grok 4 model with a 1M-token context window.
- API model ID: grok-4.3
- Price: $1.25 input / $2.50 output per 1M tokens; cached input $0.20
- Long-context rates: $2.50 in / $5.00 out above 200K prompt tokens
- Batch discount: —
- Context window: 1M tokens; max output: —
- Requests with 200K or more prompt tokens are billed at $2.50 input / $0.40 cached / $5 output for all tokens.
- Official pricing: https://docs.x.ai/docs/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/grok-4-3

### DeepSeek V4.1 Flash (DeepSeek)
DeepSeek’s low-cost model with thinking and non-thinking modes, tool calls and vision input.
- API model ID: deepseek-flash
- Price: $0.30 input / $1.20 output per 1M tokens; cached input $0.006
- Long-context rates: —
- Batch discount: —
- Context window: 1M tokens; max output: 384K tokens
- Prices shown are peak rates. Off-peak rates are half: peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday.
- Compatible with both OpenAI-format and Anthropic-format APIs.
- Official pricing: https://api-docs.deepseek.com/quick_start/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/deepseek-flash

### DeepSeek V4 Pro (DeepSeek)
DeepSeek’s larger V4 model with thinking and non-thinking modes and tool calls.
- API model ID: deepseek-v4-pro
- Price: $1.32 input / $3.96 output per 1M tokens; cached input $0.044
- Long-context rates: —
- Batch discount: —
- Context window: 1M tokens; max output: 384K tokens
- Epoch Capabilities Index: 155.5
- Prices shown are peak rates. Off-peak rates are half: peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday.
- Vision input is not supported.
- Official pricing: https://api-docs.deepseek.com/quick_start/pricing (verified September 14, 2026)
- Page: https://polarison.com/models/deepseek-v4-pro

## Methodology
- Prices are copied from each provider’s official pricing page; each model records the date it was last verified.
- Cost per request = (fresh input × input rate + cached input × cached rate + output × output rate) ÷ 1,000,000.
- Where a provider publishes a long-context threshold, prompts above it use the higher rate.
- Capability scores are the best result Epoch AI recorded for each exact model; models without an exact match stay unrated.
- Two models whose ECI confidence intervals overlap are treated as roughly tied. Benchmark gaps under 2 points are treated as even.
- A model is “best value” when no cheaper tracked model (by blended price) has a higher ECI.
- Full methodology: https://polarison.com/methodology
