Icon for Aurora Inference

Aurora Inference

OpenAI-compatible API for open-weight models, with in-region deployment on reserved capacity

Aurora Inference serves open-weight models, including DeepSeek V4, GLM 5.2 and Kimi K3, through a single OpenAI-compatible endpoint. Metered per-token access has no commitment or minimum, and reserved capacity is available for steady workloads.

On reserved capacity, prompts, outputs and weights stay in the chosen region, with deployments across North America, Europe, the Nordics, the Middle East and APAC. Cached input is priced separately, from $0.006/1M tokens on DeepSeek V4 Flash. It is part of Aurora Infra, which also sells GPU clusters and storage.

Pricing: Per token usage

Hosting Cloud
Pricing From $0.20/1M input tokens
Screenshot of Aurora Inference webpage

Work on Aurora Inference? Feature it at the top of Inference APIs.

Is your product missing?

Add it here →