≫ Home / LLM API pricing / gpt-oss-120b

gpt-oss-120b API pricing

9 providers publish a per-token price for gpt-oss-120b (OpenAI, open weights).

OpenAI's open-weight reasoning model, released in August 2025 under Apache 2.0. It was built to fit on a single 80GB GPU such as an NVIDIA H100 or AMD MI300X.

Creator
OpenAI
Released
5 August 2025
License
Apache 2.0
Size
117B total, 5B active
Context
128K tokens
Input
text
Model card
Hugging Face

Facts from the model card, checked 25 Sep 2026.

Across 9 providers, the most expensive charges 6.4Γ— the cheapest on a blended basis.

Prices by provider

Per 1M tokens, on-demand, sorted cheapest first by blended cost (3:1 input to output). Off-peak and long-context rates are listed on each provider's page.

# Provider Input / 1M Cached / 1M Output / 1M Checked Notes
1 deepinfra $0.037 – $0.17 25 Sep 2026
2 novita.ai $0.05 – $0.25 25 Sep 2026
3 SiliconFlow $0.05 – $0.45 25 Sep 2026
4 Baseten $0.10 – $0.50 25 Sep 2026
5 Groq $0.15 – $0.60 25 Sep 2026
6 together.ai $0.15 – $0.60 25 Sep 2026
7 Scaleway €0.15 β‰ˆ $0.1705 – €0.60 β‰ˆ $0.682 25 Sep 2026
8 IONOS AI Model Hub €0.15 β‰ˆ $0.1705 – €0.65 β‰ˆ $0.7389 25 Sep 2026 Price on ionos.de; the US site lists $0.17 / $0.71
9 Cerebras $0.35 – $0.75 25 Sep 2026

Each price was read off the provider's own pricing page on the date shown; the date links to that page. Providers that publish no per-token price for this model are not listed. EUR prices are ranked at the ECB reference rate of 24 September 2026. The same data is open under CC-BY at /api/query/prices.

Compare every inference provider on free tiers, hosting and EU availability in the inference APIs table.

Is your product missing?

Add it here →