Alibaba Cloud Model Studio
Hosted API access to Qwen and third-party models across six global regions
Alibaba Cloud Model Studio is a hosted inference API for Qwen models (including Qwen3.8-Max and the Qwen3.7 series) alongside third-party models like DeepSeek-V4, Kimi K2.7 and GLM-5. It also covers image, video, audio and embedding models, with an OpenAI-compatible endpoint and an Anthropic-compatible endpoint.
Six regions are live: China (Beijing), Singapore, Hong Kong, Germany (Frankfurt), Japan (Tokyo) and US (Virginia), each with its own access domain and data residency. Germany and the US support region-scoped inference for teams with data locality requirements.
Pricing is pay-as-you-go per token, tiered by context length on some models. Qwen3.8-Max runs $2/$6 per million input/output tokens. New users get a free token quota (Singapore region only) valid 90 days from activation.
Pricing: Per token usage
Alibaba Cloud Model Studio prices by model
Per 1M tokens, read off Alibaba Cloud Model Studio's own pricing page on the date shown.
| Model | Input / 1M | Cached / 1M | Output / 1M | Checked | Notes |
|---|---|---|---|---|---|
| DeepSeek V4 Pro (off peak) | $0.66 | โ | $1.98 | 25 Sep 2026 | International region, idle hours |
| DeepSeek V4 Pro | $1.32 | โ | $3.96 | 25 Sep 2026 | International region, deepseek-v4-pro-0813, busy hours; idle hours are half price |
| DeepSeek V4.1 Flash (off peak) | $0.15 | โ | $0.60 | 25 Sep 2026 | International region, idle hours |
| DeepSeek V4.1 Flash | $0.30 | โ | $1.20 | 25 Sep 2026 | International region, busy hours; idle hours are half price |
| GLM 5.3 | $1.40 | โ | $4.40 | 25 Sep 2026 | International region |
| Qwen3.8 2.4T A95B | $2.00 | โ | $6.00 | 25 Sep 2026 | International region |
Alibaba Cloud Model Studio Alternatives
Explore 115 products in the Inference APIs category. View all Alibaba Cloud Model Studio alternatives.
CheapestInference
Flat-rate unlimited inference on open-weight models, sold in daily 8-hour windows
DeepInfra
Run the top AI models using a simple API, pay per use. Low cost, scalable and production ready infrastructure.
Varion
OpenAI-compatible proxy that trims LLM token usage before requests reach your provider
LLM Tech
EU inference provider serving Qwen3.8-27B from dedicated Helsinki GPUs it rents and operates itself, with zero data r...
Work on Alibaba Cloud Model Studio? Feature it at the top of Inference APIs.
Is your product missing?