Baseten
AI inference platform for deploying and serving ML models with autoscaling and optimized infrastructure
Baseten is an inference platform for deploying ML models as production API endpoints. It provides an inference runtime with custom kernels and speculation engines, plus infrastructure that handles request routing, autoscaling, and multi-cloud capacity management. Bring any model, from open-source LLMs to fine-tuned checkpoints, and get a scalable API. Supports cloud, self-hosted, and hybrid deployments. Pricing is usage-based with no platform fee on the Startup plan, and new accounts receive $30 in free credits.
Pricing: Usage-based
Baseten prices by model
Per 1M tokens, read off Baseten's own pricing page on the date shown.
| Model | Input / 1M | Cached / 1M | Output / 1M | Checked | Notes |
|---|---|---|---|---|---|
| DeepSeek V4 Pro | $1.32 | $0.132 | $3.96 | 25 Sep 2026 | Listed as DeepSeek V4 Pro 0813 |
| DeepSeek V4.1 Flash | $0.30 | $0.007 | $1.20 | 25 Sep 2026 | |
| GLM 5.3 | $1.40 | $0.14 | $4.40 | 25 Sep 2026 | |
| GLM 5.3 Flash | $0.15 | $0.03 | $0.50 | 25 Sep 2026 | |
| Kimi K3 | $3.00 | $0.30 | $15.00 | 25 Sep 2026 | |
| gpt-oss-120b | $0.10 | โ | $0.50 | 25 Sep 2026 |
Baseten Alternatives
Explore 104 products in the Inference APIs category. View all Baseten alternatives.
BentoML
BentoML is the platform for software engineers to build AI products.
Modal
Run generative AI models, large-scale batch jobs, job queues, and much more.
Replicate
Run and fine-tune open-source models. Deploy custom models at scale. All with one line of code.
Lambda
GPU cloud for AI training and inference with on-demand and cluster options
CoreWeave
GPU cloud infrastructure built for large-scale AI training and inference workloads
Compare
Work on Baseten? Feature it at the top of Inference APIs.
Is your product missing?