≫ Home / LLM API pricing

LLM API pricing by model

The same model often costs very different amounts depending on who serves it, and the cheapest provider changes from model to model. Every price here was read off the provider's own pricing page on the date shown on the model's page. Open a model to compare its providers.

Model Released License Size Providers Lowest blended price Input / 1M Output / 1M
DeepSeek V4 Pro DeepSeek Aug 2026 MIT 1.6T 10 Aurora Inference (Listed as deepseek/deepseek-v4-pro) $1.00 $2.00
DeepSeek V4.1 Flash DeepSeek Sep 2026 MIT 552B 9 deepinfra (Promotional price, 30% off list $0.20 / $0.60) $0.14 $0.42
gpt-oss-120b OpenAI Aug 2025 Apache 2.0 117B 9 deepinfra $0.037 $0.17
GLM 5.3 Z.ai Aug 2026 GLM-5.3 License – 8 deepinfra (Promotional price, 38% off list $0.90 / $4.00) $0.5625 $2.50
GLM 5.3 Flash Z.ai Aug 2026 MIT 320B 7 deepinfra (Promotional price, 50% off list $0.15 / $0.50) $0.075 $0.25
Kimi K3 Moonshot AI – Kimi K3 License 2.8T 6 SiliconFlow $2.70 $13.50
Llama 3.3 70B Instruct Meta Dec 2024 Llama 3.3 Community License 70B 6 novita.ai $0.135 $0.40
Qwen3.8 2.4T A95B Alibaba (Qwen) Aug 2026 Qwen3.8-Max License 2.4T 5 deepinfra $2.00 $6.00

Fewer than 5 verified providers so far

  • gpt-oss-20b: 4 providers, lowest deepinfra at $0.03 input and $0.14 output per 1M
  • MiniMax M3: 2 providers, lowest deepinfra at $0.28 input and $1.10 output per 1M
  • Qwen3.8 27B: 1 provider, lowest LLM Tech at $0.25 input and $2.09 output per 1M

How these prices are collected

Candidate prices come from OpenRouter's public endpoint list. None is published until someone reads the same price on the provider's own pricing page; that page and the date are kept on every row. Where the direct price differs from OpenRouter's, the direct price is the one shown. Promotions are listed at the price charged today, with the list price in a note. EUR prices stay in EUR and rank by their value at the ECB reference rate.

The full dataset is open under CC-BY: CSV or JSON (docs). For free tiers, hosting and EU availability across every provider, see the inference APIs comparison.

Is your product missing?

Add it here →