Icon for Baseten

Baseten

Free Trial

AI inference platform for deploying and serving ML models with autoscaling and optimized infrastructure

Baseten is an inference platform for deploying ML models as production API endpoints. It provides an inference runtime with custom kernels and speculation engines, plus infrastructure that handles request routing, autoscaling, and multi-cloud capacity management. Bring any model, from open-source LLMs to fine-tuned checkpoints, and get a scalable API. Supports cloud, self-hosted, and hybrid deployments. Pricing is usage-based with no platform fee on the Startup plan, and new accounts receive $30 in free credits.

Pricing: Usage-based

Hosting Cloud + Self-hosted
Pricing Usage Based, ~$0.63/hr (T4 GPU)
HQ ๐Ÿ‡บ๐Ÿ‡ธ United States
Founded 2019
GitHub 1,155 stars
Screenshot of Baseten webpage

Baseten prices by model

Per 1M tokens, read off Baseten's own pricing page on the date shown.

Model Input / 1M Cached / 1M Output / 1M Checked Notes
DeepSeek V4 Pro $1.32 $0.132 $3.96 25 Sep 2026 Listed as DeepSeek V4 Pro 0813
DeepSeek V4.1 Flash $0.30 $0.007 $1.20 25 Sep 2026
GLM 5.3 $1.40 $0.14 $4.40 25 Sep 2026
GLM 5.3 Flash $0.15 $0.03 $0.50 25 Sep 2026
Kimi K3 $3.00 $0.30 $15.00 25 Sep 2026
gpt-oss-120b $0.10 โ€“ $0.50 25 Sep 2026

Work on Baseten? Feature it at the top of Inference APIs.

Is your product missing?

Add it here →