OpenAI · open weights
gpt-oss-120b API pricing by provider
gpt-oss-120b is available from 16 providers. The cheapest is AkashML at $0.03 input / $0.17 output per 1M tokens (BF16). The priciest host, Cerebras, charges 6.9x the cheapest.
Prices from the OpenRouter API, retrieved September 15, 2026. Weights: openai/gpt-oss-120b.
| Provider | Precision | Input / 1M | Output / 1M | Document Q&A / month | Context | Max output |
|---|---|---|---|---|---|---|
| AkashML | bf16 | $0.03 | $0.17 | $17.10 | 131K | 118K |
| CoreWeave | fp4 | $0.03 | $0.17 | $17.10 | 131K | 118K |
| DekaLLM | bf16 | $0.03 | $0.18 | $17.40 | 131K | 118K |
| DeepInfra | bf16 | $0.037 | $0.17 | $19.90 | 131K | 118K |
| Crusoe | bf16 | $0.05 | $0.25 | $27.50 | 131K | 118K |
| DigitalOcean | — | $0.06 | $0.42 | $36.60 | 128K | 4K |
| — | $0.09 | $0.36 | $46.80 | 131K | 118K | |
| Mancer 2 | fp8 | $0.055 | $0.50 | $37.00 | 131K | 118K |
| BaseTen | fp4 | $0.10 | $0.50 | $55.00 | 128K | 115K |
| Amazon Bedrock | — | $0.15 | $0.60 | $78.00 | 131K | 118K |
| Nebius | fp4 | $0.15 | $0.60 | $78.00 | 131K | 118K |
| DeepInfra | bf16 | $0.15 | $0.60 | $78.00 | 131K | 16K |
| SiliconFlow | fp8 | $0.15 | $0.60 | $78.00 | 131K | 8K |
| Groq | — | $0.15 | $0.60 | $78.00 | 131K | 66K |
| Parasail | fp4 | $0.10 | $0.75 | $62.50 | 131K | 118K |
| SambaNova | — | $0.14 | $0.95 | $84.50 | 131K | 118K |
| DeepInfra | fp8 | $0.20 | $0.95 | $108.50 | 131K | 118K |
| Cerebras | fp16 | $0.35 | $0.75 | $162.50 | 131K | 41K |
Sorted by blended price (3 input : 1 output). Document Q&A assumes 50,000 questions a month with 8,000 input and 600 output tokens each. Precision is reported by the provider; “—” means not disclosed.
Frequently asked questions
- What is the cheapest gpt-oss-120b API provider?
- AkashML is the cheapest healthy provider we track at $0.03 per 1M input tokens and $0.17 per 1M output tokens.
- How much does gpt-oss-120b cost?
- Across 16 providers, input prices range from $0.03 to $0.35 per 1M tokens and output prices from $0.17 to $0.95.
- What is gpt-oss-120b’s context window?
- gpt-oss-120b supports up to 131K tokens, but some providers serve a smaller context window — check the table before choosing.
Other open-weight models
- DeepSeek V4.1 Flashfrom $0.15 / $0.60
- DeepSeek V4 Pro 0813from $0.96 / $2.88
- Kimi K3from $2.10 / $10.95
- GLM 5.3from $0.91 / $2.86
- GLM 5.3 Flashfrom $0.075 / $0.25
- Qwen3.8 27Bfrom $0.15 / $2.00
- MiniMax M3from $0.23 / $0.96
- gpt-oss-20bfrom $0.02 / $0.10
- Llama 4 Maverickfrom $0.1875 / $0.6525
- Llama 3.3 70B Instructfrom $0.10 / $0.32
- Gemma 4 31Bfrom $0.09 / $0.34
- Mistral Small 4from $0.15 / $0.60
- Nemotron 3 Superfrom $0.085 / $0.40