≫ Home / LLM API pricing / Qwen3.8 2.4T A95B

Qwen3.8 2.4T A95B API pricing

5 providers publish a per-token price for Qwen3.8 2.4T A95B (Alibaba (Qwen), open weights).

The open-weight base of Alibaba's hosted Qwen3.8-Max. It is the first Max-class Qwen model with public weights. Context is 262K natively and extends to about 1M. Check the license: a custom Qwen3.8-Max license, not Apache 2.0 like Qwen3.8-27B.

Creator
Alibaba (Qwen)
Released
12 August 2026
License
Qwen3.8-Max License
Size
2.4T total, 95B active
Context
256K tokens
Input
text
Model card
Hugging Face

Facts from the model card, checked 25 Sep 2026.

Across 5 providers, prices sit within 0% of each other on a blended basis.

Prices by provider

Per 1M tokens, on-demand, sorted cheapest first by blended cost (3:1 input to output). Off-peak and long-context rates are listed on each provider's page.

# Provider Input / 1M Cached / 1M Output / 1M Checked Notes
1 deepinfra $2.00 $0.20 $6.00 25 Sep 2026
2 together.ai $2.00 $0.25 $6.00 25 Sep 2026
3 novita.ai $2.00 $0.25 $6.00 25 Sep 2026
4 SiliconFlow $2.00 $0.25 $6.00 25 Sep 2026
5 Alibaba Cloud Model Studio $2.00 – $6.00 25 Sep 2026 International region

Each price was read off the provider's own pricing page on the date shown; the date links to that page. Providers that publish no per-token price for this model are not listed. The same data is open under CC-BY at /api/query/prices.

Compare every inference provider on free tiers, hosting and EU availability in the inference APIs table.

Is your product missing?

Add it here →