≫ Home / LLM API pricing / DeepSeek V4.1 Flash

DeepSeek V4.1 Flash API pricing

9 providers publish a per-token price for DeepSeek V4.1 Flash (DeepSeek, open weights).

A mixture-of-experts model from DeepSeek that reads images as well as text, with a context of up to 1M tokens. The older V4 Flash and V4 Flash Vision names on DeepSeek's own API now route to it; existing integrations keep working unchanged.

Creator
DeepSeek
Released
10 September 2026
License
MIT
Size
552B total
Context
1M tokens
Input
text, image
Model card
Hugging Face

Facts from the model card, checked 25 Sep 2026.

Across 9 providers, the most expensive charges 2.5× the cheapest on a blended basis. 1 provider is on a limited-time promotion.

Prices by provider

Per 1M tokens, on-demand, sorted cheapest first by blended cost (3:1 input to output). Off-peak and long-context rates are listed on each provider's page.

# Provider Input / 1M Cached / 1M Output / 1M Checked Notes
1 deepinfra $0.14 $0.0042 $0.42 25 Sep 2026 Promotional price, 30% off list $0.20 / $0.60
2 SiliconFlow $0.15 $0.003 $0.60 25 Sep 2026
3 Aurora Inference $0.20 $0.006 $0.60 25 Sep 2026 Listed as deepseek/deepseek-v4-flash
4 DeepSeek $0.30 $0.006 $1.20 25 Sep 2026 Peak rate; off-peak hours are half price
5 together.ai $0.30 $0.006 $1.20 25 Sep 2026
6 novita.ai $0.30 $0.006 $1.20 25 Sep 2026
7 Alibaba Cloud Model Studio $0.30 – $1.20 25 Sep 2026 International region, busy hours; idle hours are half price
8 Baseten $0.30 $0.007 $1.20 25 Sep 2026
9 fireworks.ai $0.30 $0.006 $1.20 25 Sep 2026 Standard tier

Each price was read off the provider's own pricing page on the date shown; the date links to that page. Providers that publish no per-token price for this model are not listed. The same data is open under CC-BY at /api/query/prices.

Compare every inference provider on free tiers, hosting and EU availability in the inference APIs table.

Is your product missing?

Add it here →