SiliconFlow
OpenAI-compatible API serving 200+ open-source LLM and multimodal models
SiliconFlow is an inference platform that serves open-source LLMs alongside image, video, and audio models through a single OpenAI-compatible API. It hosts 200+ models, including the DeepSeek, Qwen, and Kimi families, with per-token usage pricing and serverless deployment.
It also offers reserved GPU options for predictable billing. Developers use it as a drop-in alternative to other hosted inference APIs, switching by changing the base URL and key.
Pricing: Per token usage
SiliconFlow prices by model
Per 1M tokens, read off SiliconFlow's own pricing page on the date shown.
| Model | Input / 1M | Cached / 1M | Output / 1M | Checked | Notes |
|---|---|---|---|---|---|
| DeepSeek V4 Pro | $1.32 | $0.044 | $3.96 | 25 Sep 2026 | Listed as DeepSeek-V4-Pro-0813 |
| DeepSeek V4.1 Flash | $0.15 | $0.003 | $0.60 | 25 Sep 2026 | |
| GLM 5.3 | $1.40 | $0.26 | $4.40 | 25 Sep 2026 | |
| GLM 5.3 Flash | $0.15 | $0.03 | $0.50 | 25 Sep 2026 | |
| Kimi K3 | $2.70 | $0.27 | $13.50 | 25 Sep 2026 | |
| Qwen3.8 2.4T A95B | $2.00 | $0.25 | $6.00 | 25 Sep 2026 | |
| gpt-oss-120b | $0.05 | – | $0.45 | 25 Sep 2026 | |
| gpt-oss-20b | $0.04 | – | $0.18 | 25 Sep 2026 |
SiliconFlow Alternatives
Explore 120 products in the Inference APIs category. View all SiliconFlow alternatives.
CheapestInference
Flat-rate unlimited inference on open-weight models, sold in daily 8-hour windows
LLM Tech
EU inference provider serving Qwen3.8-27B from dedicated Helsinki GPUs it rents and operates itself, with zero data r...
Voyage AI
Embedding and reranker models for RAG retrieval quality, from MongoDB
Work on SiliconFlow? Feature it at the top of Inference APIs.
Is your product missing?