Polarison

AI model benchmarks vs. API price

The most capable model we track is GPT-6 Astra (Epoch Capabilities Index 166.3). For the best capability per dollar, look at GPT-5.6 Luna, Gemini 3.8 Flash, GPT-5.6 Terra, GPT-5.6 Sol, Claude Opus 5, GPT-6 Astra: no cheaper model beats any of them.

Benchmark data from Epoch AI, retrieved September 15, 2026.

Capability vs. price

Epoch Capabilities Index (higher is more capable) against blended API price per 1M tokens. Hover or tab to a dot for details.

Best value: no cheaper model scores higherOther modelsConfidence interval
135140145150155160165170175$0.10$0.30$1$3$10$30Blended price per 1M tokens (3 input : 1 output, log scale)Epoch Capabilities IndexGPT-5.6 LunaGemini 3.8 FlashGPT-5.6 TerraGPT-5.6 SolClaude Opus 5GPT-6 Astra

Best model for each job

Leaderboard

Sorted by Epoch Capabilities Index. Prices are blended USD per 1M tokens (3 input : 1 output).

#ModelECIPriceGPQA DiamondFrontierMath (Tiers 1–3)SimpleQA VerifiedARC-AGI-2DeepSWEAPEX-Agents
1GPT-6 AstraOpenAIBest value166.3(163172)$20.0094.4%93.7%75.6%95.0%74.1%46.7%
2Claude Fable 5.1Anthropic164.5(161168)$20.0090.2%70.8%90.0%47.4%
3Claude Opus 5AnthropicBest value162.3(160166)$10.0091.8%85.6%59.9%90.4%73.6%43.5%
4GPT-5.6 SolOpenAIBest value161.8(160165)$8.0091.3%89.1%69.7%92.5%72.7%40.0%
5GPT-5.6 TerraOpenAIBest value159.1(157162)$4.5091.1%86.0%43.2%83.9%69.6%
6Gemini 3.8 FlashGoogleBest value156.5(155162)$1.5093.9%68.4%69.7%73.8%
7Grok 4.6xAI156.3(155159)$3.0092.0%66.0%49.3%67.1%67.5%41.2%
8GPT-5.6 LunaOpenAIBest value156.3(154159)$0.4588.8%82.1%41.0%59.5%67.2%
9Claude Sonnet 5Anthropic156.2(154159)$4.0087.4%65.6%33.7%53.8%32.5%
10DeepSeek V4 ProDeepSeek155.5(154158)$1.9888.9%64.6%52.9%61.3%
11Gemini 3.1 ProGoogle155.0(153158)$4.5092.6%59.6%73.5%77.1%11.7%33.5%
12Gemini 3.5 FlashGoogle154.7(153157)$3.37590.4%62.8%66.2%72.1%37.4%
13Gemini 2.5 ProGoogle145.3(144147)$3.43880.4%24.6%4.9%6.6%
14Gemini 3.5 Flash-LiteGoogle145.1(143147)$0.8577.8%26.0%10.3%
15Claude Haiku 4.5Anthropic142.4(140144)$2.0061.6%13.2%4.0%8.9%
16Gemini 2.5 FlashGoogle140.5(138142)$0.851.8%

Not yet rated by Epoch AI: Grok 4.3, DeepSeek V4.1 Flash.

How to read these scores

The Epoch Capabilities Index combines many benchmarks into one general-capability scale. The numbers in brackets are its confidence interval: when two models’ intervals overlap, treat them as roughly equal rather than declaring a winner.

Each benchmark score is the best result Epoch AI recorded for that model, usually at its highest reasoning setting. Higher reasoning settings produce more output tokens, so a model can cost more in practice than its per-token price suggests.

The benchmarks

  • GPQA Diamond (science reasoning): Graduate-level biology, chemistry and physics questions designed to be hard to look up.
  • FrontierMath (Tiers 1–3) (advanced math): Original, unpublished math problems ranging from advanced undergraduate to research level.
  • SimpleQA Verified (factual accuracy): Short fact-seeking questions that reward correct answers over confident guesses.
  • ARC-AGI-2 (abstract reasoning): Visual pattern puzzles that are easy for people but hard for AI systems.
  • DeepSWE (coding): Agentic software engineering tasks that require changing real codebases.
  • APEX-Agents (agentic work): Long, multi-step professional tasks completed by an AI agent.

A dash means Epoch AI has not published a score for that model on that benchmark. Benchmarks measure narrow skills; test candidates on your own tasks before committing. To turn these numbers into a monthly bill, use the cost calculator or compare two models.

Data: Epoch AI, “Capabilities & Benchmarking”, epoch.ai/benchmarks, licensed CC BY 4.0. Prices from each provider’s official pricing page.

Frequently asked questions

What is the most capable AI model right now?
Among the models we track, GPT-6 Astra from OpenAI has the highest Epoch Capabilities Index score (166.3), followed by Claude Fable 5.1 (164.5) and Claude Opus 5 (162.3). Scores this close can fall within each other’s confidence intervals.
Which AI model is best for coding?
On the DeepSWE agentic coding benchmark, GPT-6 Astra scores highest among tracked models at 74.1%.
What is the best value AI model?
These models offer the most capability for their price, because no cheaper model we track scores higher: GPT-5.6 Luna ($0.45 per 1M tokens), Gemini 3.8 Flash ($1.50 per 1M tokens), GPT-5.6 Terra ($4.50 per 1M tokens), GPT-5.6 Sol ($8.00 per 1M tokens), Claude Opus 5 ($10.00 per 1M tokens), GPT-6 Astra ($20.00 per 1M tokens).
What is the Epoch Capabilities Index?
The Epoch Capabilities Index (ECI), published by Epoch AI, combines results from many AI benchmarks into a single general-capability scale, so models can be compared even after individual benchmarks saturate. Higher is better.