≫ Home / LLM API pricing / GLM 5.3 Flash

GLM 5.3 Flash API pricing

7 providers publish a per-token price for GLM 5.3 Flash (Z.ai, open weights).

The first GLM 5 model to take images as input. Only 18B of its 320B parameters are active per token. Z.ai prices it at about a tenth of GLM 5.3 and the providers here follow: $0.15 / $0.50 per 1M against $1.40 / $4.40.

Creator
Z.ai
Released
26 August 2026
License
MIT
Size
320B total, 18B active
Input
text, image
Model card
Hugging Face

Facts from the model card, checked 25 Sep 2026.

Across 7 providers, the most expensive charges 2.0× the cheapest on a blended basis. 1 provider is on a limited-time promotion.

Prices by provider

Per 1M tokens, on-demand, sorted cheapest first by blended cost (3:1 input to output). Off-peak and long-context rates are listed on each provider's page.

# Provider Input / 1M Cached / 1M Output / 1M Checked Notes
1 deepinfra $0.075 $0.015 $0.25 25 Sep 2026 Promotional price, 50% off list $0.15 / $0.50
2 together.ai $0.15 $0.03 $0.50 25 Sep 2026
3 novita.ai $0.15 $0.03 $0.50 25 Sep 2026
4 SiliconFlow $0.15 $0.03 $0.50 25 Sep 2026
5 Baseten $0.15 $0.03 $0.50 25 Sep 2026
6 fireworks.ai $0.15 $0.03 $0.50 25 Sep 2026 Standard tier
7 Cloudflare Workers AI $0.15 $0.03 $0.50 25 Sep 2026

Each price was read off the provider's own pricing page on the date shown; the date links to that page. Providers that publish no per-token price for this model are not listed. The same data is open under CC-BY at /api/query/prices.

Compare every inference provider on free tiers, hosting and EU availability in the inference APIs table.

Is your product missing?

Add it here →