Z.ai · open weights
GLM 5.3 API pricing by provider
GLM 5.3 is available from 23 providers. The cheapest is Phala at $0.91 input / $2.86 output per 1M tokens. Z.ai’s own API charges $1.40 / $4.40, 35% more than the cheapest host. The priciest host, BaseTen, charges 2.3x the cheapest.
Prices from the OpenRouter API, retrieved September 15, 2026. Weights: zai-org/GLM-5.3.
| Provider | Precision | Input / 1M | Output / 1M | Document Q&A / month | Context | Max output |
|---|---|---|---|---|---|---|
| Phala | — | $0.91 | $2.86 | $449.80 | 1.05M | 131K |
| Reka | fp8 | $0.8775 | $2.97 | $440.10 | 262K | 236K |
| DeepInfra | fp4 | $0.90 | $3.00 | $450.00 | 1.05M | 131K |
| Morph | fp8 | $0.90 | $3.069 | $452.07 | 1.05M | 944K |
| Novita | fp8 | $1.092 | $3.432 | $539.76 | 1.05M | 131K |
| GMICloud | fp8 | $1.12 | $3.52 | $553.60 | 1.05M | 944K |
| Wafer | — | $0.95 | $4.40 | $512.00 | 1.05M | 944K |
| Inceptron | fp4 | $1.07 | $4.092 | $550.72 | 1.05M | 944K |
| Decart | fp4 | $1.19 | $3.74 | $588.20 | 1.05M | 944K |
| Sail Research | fp8 | $1.257 | $3.951 | $621.42 | 1.05M | 131K |
| DigitalOcean | — | $1.26 | $3.96 | $622.80 | 1.05M | 128K |
| Friendli | — | $1.26 | $3.96 | $622.80 | 1.05M | 944K |
| AkashML | fp8 | $1.30 | $4.40 | $652.00 | 1.05M | 944K |
| Baidu | fp8 | $1.40 | $4.40 | $692.00 | 1.05M | 131K |
| SiliconFlow | fp8 | $1.40 | $4.40 | $692.00 | 1.05M | 262K |
| Together | — | $1.40 | $4.40 | $692.00 | 1.05M | 944K |
| Parasail | fp8 | $1.40 | $4.40 | $692.00 | 1.05M | 944K |
| Modal | — | $1.40 | $4.40 | $692.00 | 1.05M | 944K |
| BaseTen | fp4 | $1.40 | $4.40 | $692.00 | 1.05M | 262K |
| Fireworks | — | $1.40 | $4.40 | $692.00 | 1.05M | 944K |
| Cloudflare | — | $1.40 | $4.40 | $692.00 | 1.31M | 1.18M |
| AtlasCloud | fp8 | $1.40 | $4.40 | $692.00 | 262K | 131K |
| Z.AIcreator | fp8 | $1.40 | $4.40 | $692.00 | 1.05M | 131K |
| BaseTen | fp8 | $2.10 | $6.60 | $1,038 | 1.05M | 262K |
Sorted by blended price (3 input : 1 output). Document Q&A assumes 50,000 questions a month with 8,000 input and 600 output tokens each. Precision is reported by the provider; “—” means not disclosed.
Frequently asked questions
- What is the cheapest GLM 5.3 API provider?
- Phala is the cheapest healthy provider we track at $0.91 per 1M input tokens and $2.86 per 1M output tokens.
- How much does GLM 5.3 cost?
- Across 23 providers, input prices range from $0.8775 to $2.10 per 1M tokens and output prices from $2.86 to $6.60.
- What is GLM 5.3’s context window?
- GLM 5.3 supports up to 1.31M tokens, but some providers serve a smaller context window — check the table before choosing.
Other open-weight models
- DeepSeek V4.1 Flashfrom $0.15 / $0.60
- DeepSeek V4 Pro 0813from $0.96 / $2.88
- Kimi K3from $2.10 / $10.95
- GLM 5.3 Flashfrom $0.075 / $0.25
- Qwen3.8 27Bfrom $0.15 / $2.00
- MiniMax M3from $0.23 / $0.96
- gpt-oss-120bfrom $0.03 / $0.17
- gpt-oss-20bfrom $0.02 / $0.10
- Llama 4 Maverickfrom $0.1875 / $0.6525
- Llama 3.3 70B Instructfrom $0.10 / $0.32
- Gemma 4 31Bfrom $0.09 / $0.34
- Mistral Small 4from $0.15 / $0.60
- Nemotron 3 Superfrom $0.085 / $0.40