Polarison

Open-weight model API prices by provider

The same open-weight model can cost very different amounts depending on who hosts it. gpt-oss-120b runs on 16 providers, from $0.03 to $0.35 per 1M input tokens. That 6.9x gap between the priciest and cheapest host is the widest we track.

Provider prices from the OpenRouter API, retrieved September 15, 2026. Only endpoints reporting healthy status are included.

ModelProvidersCheapest hostCheapest (in / out)Creator’s own APIPrice spreadContext
gpt-oss-20bOpenAI13Darkbloomfp8$0.02 / $0.103.3x131K
gpt-oss-120bOpenAI16AkashMLbf16$0.03 / $0.176.9x131K
GLM 5.3 FlashZ.ai24DeepInfrafp4$0.075 / $0.25$0.15 / $0.506.0x1.31M
Gemma 4 31BGoogle11DeepInfrafp4$0.09 / $0.345.3x262K
Llama 3.3 70B InstructMeta10DeepInfrafp8$0.10 / $0.326.7x131K
Nemotron 3 SuperNVIDIA2DeepInfrabf16$0.085 / $0.401.1x262K
DeepSeek V4.1 FlashDeepSeek16Relacefp4$0.15 / $0.60$0.30 / $1.202.5x1.05M
Mistral Small 4Mistral AI2Mistral$0.15 / $0.60$0.15 / $0.601.3x262K
Llama 4 MaverickMeta5DigitalOcean$0.1875 / $0.65251.8x1.05M
MiniMax M3MiniMax13CoreWeavefp4$0.23 / $0.96$0.30 / $1.203.2x1.05M
Qwen3.8 27BAlibaba (Qwen)15Darkbloomfp4$0.15 / $2.00$0.425 / $2.551.9x1M
GLM 5.3Z.ai23Phala$0.91 / $2.86$1.40 / $4.402.3x1.31M
DeepSeek V4 Pro 0813DeepSeek18Ionstream$0.96 / $2.88$1.32 / $3.961.7x1.05M
Kimi K3Moonshot AI17InferenceNet$2.10 / $10.95$3.00 / $15.002.3x1.05M

USD per 1M tokens. Sorted by the cheapest host’s blended price (3 input : 1 output). Price spread compares the priciest and cheapest healthy host.

How to pick a provider

  • Price is only one axis. Hosts differ in speed, uptime, rate limits and supported features such as tool calling. Our tables show list prices only.
  • Check the precision. Many of the cheapest endpoints serve FP8 or FP4 weights. That is usually fine for chat and extraction but can matter for math, coding and long reasoning.
  • Compare with closed models. Some open-weight models now rival closed ones on benchmarks. See how they stack up on the benchmarks page, and compare first-party APIs on our API comparison pages.

Frequently asked questions

What is the cheapest open-weight LLM API?
Among the models we track, gpt-oss-20b is cheapest to run: Darkbloom charges $0.02 per 1M input tokens and $0.10 per 1M output tokens.
Why do providers charge different prices for the same model?
Open-weight models can be hosted by anyone. Providers run them on different hardware, at different precision (quantization), with different speed and reliability guarantees, and set their own prices.
Are quantized models worse?
Serving weights at lower precision, such as FP8 or FP4, cuts hosting costs and can reduce output quality, especially on harder tasks. The cheapest endpoints are often quantized, so test quality on your own prompts before switching.