Qwen3.8 2.4T A95B API pricing
5 providers publish a per-token price for Qwen3.8 2.4T A95B (Alibaba (Qwen), open weights).
The open-weight base of Alibaba's hosted Qwen3.8-Max. It is the first Max-class Qwen model with public weights. Context is 262K natively and extends to about 1M. Check the license: a custom Qwen3.8-Max license, not Apache 2.0 like Qwen3.8-27B.
- Creator
- Alibaba (Qwen)
- Released
- 12 August 2026
- License
- Qwen3.8-Max License
- Size
- 2.4T total, 95B active
- Context
- 256K tokens
- Input
- text
- Model card
- Hugging Face
Facts from the model card, checked 25 Sep 2026.
Across 5 providers, prices sit within 0% of each other on a blended basis.
Prices by provider
Per 1M tokens, on-demand, sorted cheapest first by blended cost (3:1 input to output). Off-peak and long-context rates are listed on each provider's page.
| # | Provider | Input / 1M | Cached / 1M | Output / 1M | Checked | Notes |
|---|---|---|---|---|---|---|
| 1 | deepinfra | $2.00 | $0.20 | $6.00 | 25 Sep 2026 | |
| 2 | together.ai | $2.00 | $0.25 | $6.00 | 25 Sep 2026 | |
| 3 | novita.ai | $2.00 | $0.25 | $6.00 | 25 Sep 2026 | |
| 4 | SiliconFlow | $2.00 | $0.25 | $6.00 | 25 Sep 2026 | |
| 5 | Alibaba Cloud Model Studio | $2.00 | – | $6.00 | 25 Sep 2026 | International region |
Each price was read off the provider's own pricing page on the date shown; the date links to that page. Providers that publish no per-token price for this model are not listed. The same data is open under CC-BY at /api/query/prices.
Compare every inference provider on free tiers, hosting and EU availability in the inference APIs table.
Is your product missing?