DeepSeek
Cost-effective inference API with OpenAI-compatible endpoints and open-weight models
DeepSeek offers an inference API for its DeepSeek-V3 and reasoning models with OpenAI-compatible endpoints. Known for strong performance at significantly lower cost than competitors. Models support 128K context windows with both standard chat and chain-of-thought reasoning modes. The underlying models are open-weight, available on Hugging Face.
Pricing: Per token usage
DeepSeek prices by model
Per 1M tokens, read off DeepSeek's own pricing page on the date shown.
| Model | Input / 1M | Cached / 1M | Output / 1M | Checked | Notes |
|---|---|---|---|---|---|
| DeepSeek V4 Pro (off peak) | $0.66 | $0.022 | $1.98 | 25 Sep 2026 | Off-peak hours (most of the week) |
| DeepSeek V4 Pro | $1.32 | $0.044 | $3.96 | 25 Sep 2026 | Peak rate; off-peak hours are half price |
| DeepSeek V4.1 Flash (off peak) | $0.15 | $0.003 | $0.60 | 25 Sep 2026 | Off-peak hours (most of the week) |
| DeepSeek V4.1 Flash | $0.30 | $0.006 | $1.20 | 25 Sep 2026 | Peak rate; off-peak hours are half price |
DeepSeek Alternatives
Explore 107 products in the Inference APIs category. View all DeepSeek alternatives.
CheapestInference
Flat-rate unlimited inference on open-weight models, sold in daily 8-hour windows
LLM Tech
EU inference provider serving Qwen3.8-27B from dedicated Helsinki GPUs it rents and operates itself, with zero data r...
Cloudflare Workers AI
Run AI models at the edge on Cloudflare's global network with serverless inference
Berget AI
EU-sovereign AI inference platform with OpenAI-compatible API
Nebius
Full-stack AI cloud with GPU infrastructure for training and inference
Work on DeepSeek? Feature it at the top of Inference APIs.
Is your product missing?