Inference APIs Pricing Comparison
118 providers compared by pricing model, free tiers, hosting options, and headquarters. Last updated October 2026.
63 with free tiers · 14 open source · 22 self-hostable · 40 European
AI inference API providers are hosted services that run large language models behind an API, so you pay per token or per second instead of renting and managing GPUs yourself. The table below compares 118 of them on the axes that usually decide the choice: price per token, sustained throughput and first-token latency, the model catalog, OpenAI-compatibility, free tiers, and hosting region (40 run inside the EU). Cost-sensitive and high-throughput workloads tend to pull toward different providers, so the right pick depends on which of these matters most for your use case.
Cheapest gpt-oss-120b providers, ranked by verified price
Per-token price for one model, gpt-oss-120b, across the 15 providers that publish it, so the figures compare like for like. Sorted cheapest first by blended cost (3:1 input to output). The cheapest provider differs by model.
| # | Provider | Input / 1M | Cached / 1M | Output / 1M | Checked | Notes |
|---|---|---|---|---|---|---|
| 1 | $0.03 | – | $0.17 | 27 Sep 2026 | ||
| 2 | $0.037 | – | $0.17 | 25 Sep 2026 | ||
| 3 | €0.04 ≈ $0.0455 | – | €0.20 ≈ $0.2273 | 1 Oct 2026 | ||
| 4 | $0.05 | – | $0.25 | 25 Sep 2026 | ||
| 5 | $0.05 | – | $0.45 | 25 Sep 2026 | ||
| 6 | $0.10 | – | $0.50 | 25 Sep 2026 | ||
| 7 | $0.15 | – | $0.60 | 25 Sep 2026 | ||
| 8 | $0.15 | – | $0.60 | 25 Sep 2026 | ||
| 9 | €0.15 ≈ $0.1705 | – | €0.60 ≈ $0.682 | 25 Sep 2026 | ||
| 10 | €0.15 ≈ $0.1705 | – | €0.65 ≈ $0.7389 | 25 Sep 2026 | Price on ionos.de; the US site lists $0.17 / $0.71 | |
| 11 | €0.16 ≈ $0.1819 | – | €0.63 ≈ $0.7161 | 1 Oct 2026 | ||
| 12 | €0.16 ≈ $0.1819 | – | €0.83 ≈ $0.9435 | 1 Oct 2026 | Listed in SimpleCredits (100 SC = EUR 1); homepage advertises 20% off during the open beta | |
| 13 | $0.35 | – | $0.75 | 25 Sep 2026 | ||
| 14 | $0.35 | – | $0.75 | 27 Sep 2026 | @cf/openai/gpt-oss-120b | |
| 15 | €1.00 ≈ $1.1367 | – | €4.20 ≈ $4.7741 | 1 Oct 2026 |
Each price was read off the provider's own pricing page on the date shown; the date links to that page. Providers that publish no per-token price for this model are not listed. EUR prices are ranked at the ECB reference rate of 24 September 2026. The same data is open under CC-BY at /api/query/prices.
Prices for DeepSeek, GLM, Kimi, Llama and other models are on the LLM API pricing by model page.
| Provider | Pricing Model | Starting Price | Free Tier | Hosting | Open Source | HQ |
|---|---|---|---|---|---|---|
| Pay-per-use | From $0.06/1M tokens | — | Cloud | — | 🇩🇪 Germany | |
| Pay-per-use | Pay per prompt, from $0.10 | ✓ | Cloud | — | — | |
| Subscription | From EUR 99/month | ✓ | Cloud | — | 🇩🇪 Germany | |
| Pay-per-use | Free tier, then $0.10/1,000 requests | ✓ | Cloud | — | 🇺🇸 United States | |
| Freemium | Free | ✓ | Cloud + Self-hosted | ✓ | — | |
| Pay-per-use | Free (10 EUR credits) | ✓ | — | — | 🇩🇪 Germany | |
| Freemium | — | ✓ | Cloud + Self-hosted | — | — | |
| Subscription | 49 SEK/month | ✓ | Cloud | — | 🇸🇪 Sweden | |
| Pay-per-use | $1.24/hour | — | Cloud | — | 🇸🇪 Sweden | |
| Pay-per-use | Pay-as-you-go, EUR 10 minimum top-up | — | Cloud | — | 🇳🇱 Netherlands | |
| — | Free tier (Singapore only), then $2/1M input tokens (Qwen3.8-Max) | ✓ | — | — | — | |
| Pay-per-use | Pay-per-token | ✓ | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | $1/1M tokens | ✓ | Cloud | — | 🇺🇸 United States | |
| — | — | — | — | — | 🇺🇸 United States | |
| — | From $0.20/1M input tokens | — | Cloud | — | 🇵🇱 Poland | |
| Pay-per-use | ~$0.63/hr (T4 GPU) | ✓ | Cloud + Self-hosted | — | 🇺🇸 United States | |
| Pay-per-use | $0.15/hr (T4 GPU) | ✓ | Cloud + Self-hosted | ✓ | 🇺🇸 United States | |
| — | — | ✓ | Cloud + Self-hosted | ✓ | 🇺🇸 United States | |
| Freemium | €25/mo | ✓ | Cloud | — | 🇸🇪 Sweden | |
| — | $0.04/1M input tokens | — | Cloud | — | 🇺🇸 United States | |
| Freemium | Free tier available | ✓ | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | ~$1.10/hr (A10 GPU) | ✓ | Cloud | — | 🇺🇸 United States | |
| Subscription | $22/month (Core pool, one daily 8-hour window) | — | Cloud | — | — | |
| Pay-per-use | $0.09 per 1M input tokens (Clef-flash on Workers AI) | — | Cloud + Self-hosted | ✓ | 🇺🇸 United States | |
| Freemium | $0.011/1K neurons | ✓ | Cloud | — | 🇺🇸 United States | |
| — | — | ✓ | Cloud | — | — | |
| Freemium | $0.04/1M tokens | ✓ | Cloud + Self-hosted | — | 🇨🇦 Canada | |
| Pay-per-use | From $0.15 per 1M input tokens | — | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | $6.50/hr (GH200 GPU) | — | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | Pay-per-use + 5% gateway fee | ✓ | Cloud | — | 🇦🇹 Austria | |
| Pay-per-use | $0.02/M tokens | ✓ | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | $0.028/1M tokens (cache hit) | ✓ | Cloud | ✓ | 🇨🇳 China | |
| Subscription | Free (10K req/mo), 39 EUR/mo (Plus) | ✓ | Cloud | — | 🇳🇱 Netherlands | |
| — | — | — | — | — | — | |
| — | — | — | — | — | — | |
| Pay-per-use | $0.10/1M tokens | — | Cloud + Self-hosted | — | 🇺🇸 United States | |
| — | From EUR 0.06/1M input tokens, EUR 50 free credit on signup | ✓ | — | — | 🇸🇵🇦🇮🇳 SPAIN | |
| Pay-per-use | Pay-as-you-go, $100 free credit on signup | — | Cloud | — | 🇺🇸 United States | |
| — | — | — | Cloud | — | 🇺🇸 United States | |
| Freemium | Free | ✓ | Cloud | — | 🇺🇸 United States | |
| Subscription | 14-day free trial, then EUR 4.50/month | ✓ | Cloud | — | 🇳🇱 Netherlands | |
| Freemium | $0.05/1M tokens | ✓ | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | From $0.04 per 1M input tokens | — | Cloud | — | 🇸🇰 Slovakia | |
| Pay-per-use | From EUR 0.08/1M input tokens, prepaid balance | — | Cloud + Self-hosted | — | 🇳🇪🇹🇭🇪🇷🇱🇦🇳🇩🇸 NETHERLANDS | |
| Pay-per-use | — | — | Cloud | — | 🇦🇪 AE | |
| Pay-per-use | $0.15/hr | — | Cloud | — | 🇬🇧 United Kingdom | |
| Pay-per-use | From $0.11/1M tokens (Qwen3.5-9B) | — | Cloud | — | 🇩🇪 Germany | |
| Pay-per-use | Prepaid credits, USD 20 minimum funding | — | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | From €0.20/M input tokens (gemma-4-31B) | ✓ | Cloud + Self-hosted | — | 🇱🇺 Luxembourg | |
| Pay-per-use | $0.02/M tokens | ✓ | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | Free (10M tokens) | ✓ | Cloud | — | 🇩🇪 Germany | |
| — | — | ✓ | — | ✓ | — | |
| Free | Free (open-source) | — | Self-hosted | ✓ | — | |
| Pay-per-use | — | — | Cloud | — | 🇩🇪 Germany | |
| Pay-per-use | $0.25/1M input tokens | ✓ | Cloud | — | 🇵🇱 Poland | |
|
L
LLMBase
|
— | — | — | — | — | — |
| Freemium | $3/300 credits | ✓ | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | $0.58/GPU/hr (V100) | ✓ | Cloud | — | 🇺🇸 United States | |
| Free | Free (open-source) | — | Self-hosted | ✓ | — | |
| Freemium | Free tier, then €8/month | — | Cloud | — | 🇫🇷 France | |
| Pay-per-use | H100 from $1.30/GPU-hour (marketplace rates, Sep 2026) | — | Cloud | — | 🇰🇳 KN | |
| Pay-per-use | From EUR 49/month | — | Cloud | — | 🇩🇪 Germany | |
| Pay-per-use | — | — | Cloud | — | 🇸🇬 Singapore | |
| Pay-per-use | $1.25 / $4.25 per 1M tokens (Muse Spark) | — | — | — | 🇺🇸 United States | |
| Pay-per-use | Free tier, 100 queries/month, then from $3.60/1,000 queries | ✓ | Cloud | — | 🇬🇧 United Kingdom | |
| — | — | — | Cloud | — | — | |
| Freemium | $0.10/1M tokens | — | Cloud + Self-hosted | ✓ | 🇫🇷 France | |
| Pay-per-use | $30/mo free credits | ✓ | Cloud | — | 🇺🇸 United States | |
| Usage Based | — | — | Cloud | — | — | |
| Pay-per-use | $2.00/hr (H100) | ✓ | Cloud | — | 🇳🇱 Netherlands | |
| Pay-per-use | $0.03/M tokens | ✓ | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | $0.01/M tokens | ✓ | Cloud | — | 🇬🇧 United Kingdom | |
| Pay-per-use | $0.91/hr (L4 GPU) | ✓ | Cloud | — | 🇫🇷 France | |
| — | — | ✓ | — | — | 🇺🇸 United States | |
| Pay-per-use | $0.05/1M tokens | ✓ | Cloud | — | 🇺🇸 United States | |
| Freemium | Free (25+ free models) | ✓ | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | Free tier (BlockRun), then usage-based from $0.001/request | ✓ | Cloud | — | — | |
| Pay-per-use | Provider token rates, no markup; 3% fee on credit purchases | ✓ | — | — | 🇸🇪 Sweden | |
| — | — | — | — | — | — | |
| Pay-per-use | $0.39/hr RTX 4090 (Dedicated) | — | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | Free to prove, billed only after activation | ✓ | Cloud + Self-hosted | — | 🇬🇧 United Kingdom | |
| — | $0.004 per 1M tokens (pplx-embed-v1-0.6b) | — | Cloud | — | — | |
| — | $0.04 per 1M input tokens | ✓ | Cloud + Self-hosted | ✓ | — | |
| Contact Sales | — | — | Cloud + Self-hosted | — | 🇨🇭 Switzerland | |
| Free | Free (open source) | ✓ | Self-hosted | ✓ | — | |
| — | — | ✓ | — | — | 🇮🇹 Italy | |
| Pay-per-use | Per-second GPU billing | — | Cloud | — | 🇺🇸 United States | |
| Freemium | Free tier, then pay-as-you-go at 5% markup | ✓ | Cloud | — | — | |
| Pay-per-use | $0.06/hr | — | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | Pay-as-you-go, $2 free credit (business email) | ✓ | Cloud | — | 🇺🇸 United States | |
| Free | Free (open source) | — | Self-hosted | ✓ | — | |
| Freemium | $5 free credit | ✓ | Cloud + Self-hosted | — | 🇺🇸 United States | |
| Pay-per-use | From €0.15/M input tokens (chat models) | ✓ | Cloud | — | 🇫🇷 France | |
| Pay-per-use | Pay-as-you-go, $1 free credit | ✓ | Cloud | — | — | |
| — | — | ✓ | — | — | 🇩🇪 Germany | |
| Subscription | From EUR 15/month | — | Cloud | — | 🇮🇹 Italy | |
| Pay-per-use | $0.12 / 1M input tokens | ✓ | Cloud | — | 🇫🇷 France | |
| Pay-per-use | $0.0015/image | — | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | $0.025 per 1M input tokens, output free | — | — | — | 🇩🇪 Germany | |
| Pay-per-use | ~€2.70/GPU-hr | — | Cloud | — | 🇩🇪 Germany | |
| Pay-per-use | From $0.02/M tokens | — | Cloud | — | 🇮🇪 Ireland | |
| — | $0.15/hour | — | — | ✓ | — | |
| Pay-per-use | Pay-per-token | — | Cloud + Self-hosted | — | 🇺🇸 United States | |
| Pay-per-use | From $0.0880/1M tokens (Qwen3.5-plus) | ✓ | — | — | 🇭🇰 HK | |
| — | — | ✓ | — | — | — | |
| Pay-per-use | $0.042 per 1M input tokens | — | Cloud | — | 🇺🇸 United States | |
| — | From EUR 19/month (10M tokens) | ✓ | — | — | 🇱🇹 Lithuania | |
| Pay-per-use | ~$0.06/GPU/hr | — | Cloud | — | 🇺🇸 United States | |
| Freemium | Free tier, then pay-as-you-go at provider rates | ✓ | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | $0.14/hr | — | Cloud | — | 🇫🇮 Finland | |
| Pay-per-use | Free tier (200M tokens), then from $0.02/1M input tokens | ✓ | Cloud | — | 🇺🇸 United States | |
| Pay-per-use | From $5/month | — | Cloud | — | 🇮🇳 India | |
| Subscription | From EUR 11.99/month | — | Cloud | — | 🇳🇴 Norway | |
| Usage Based | — | — | Cloud | — | — | |
| Contact Sales | — | — | Cloud | — | 🇸🇪 Sweden | |
| Pay-per-use | $0.02/megapixel | ✓ | Cloud | — | 🇺🇸 United States | |
| Free | Free (open-source) | ✓ | Self-hosted | ✓ | 🇺🇸 United States | |
| — | — | — | Self-hosted | — | — |
Ordering is alphabetical, apart from featured listings, which are paid placement and are labelled as such. Nothing here is ranked by quality. Want to be featured?
Providers with free tiers
These inference apis providers offer free credits, free tiers, or open-source self-hosting options to get started without upfront costs.
One OpenAI-compatible API for 600+ models, with text billed at provider list ...
From: Pay per prompt, from $0.10
EU-hosted AI integration platform combining a provider gateway, document extr...
From: From EUR 99/month
LLM gateway built around spend limits enforced before a request reaches a pro...
From: Free tier, then $0.10/1,000 requests
Sovereign AI inference infrastructure for regulated EU environments, with het...
Show all 63 providers with free tiers
Swedish GPU infrastructure and LLM hosting platform with API-first deployment...
From: 49 SEK/month
Hosted API access to Qwen and third-party models across six global regions
From: Free tier (Singapore only), then $2/1M input tokens (Qwen3.8-Max)
Managed API access to foundation models on AWS with built-in fine-tuning and ...
From: Pay-per-token
Claude API for building AI applications with Opus, Sonnet, and Haiku models
From: $1/1M tokens
AI inference platform for deploying and serving ML models with autoscaling an...
From: ~$0.63/hr (T4 GPU)
Open-source serverless GPU cloud with sub-second cold starts and auto-scaling
From: $0.15/hr (T4 GPU)
BentoML is the platform for software engineers to build AI products.
Ultra-fast inference on custom wafer-scale hardware with OpenAI-compatible API
From: Free tier available
Serverless GPU infrastructure for deploying AI models with sub-5 second cold ...
From: ~$1.10/hr (A10 GPU)
Run AI models at the edge on Cloudflare's global network with serverless infe...
From: $0.011/1K neurons
Unified AI API gateway providing access to 600+ models from OpenAI, Anthropic...
Cohere’s world-class LLMs help enterprises build powerful, secure application...
From: $0.04/1M tokens
European AI inference gateway with smart routing across EU providers
From: Pay-per-use + 5% gateway fee
Run the top AI models using a simple API, pay per use. Low cost, scalable and...
From: $0.02/M tokens
Cost-effective inference API with OpenAI-compatible endpoints and open-weight...
From: $0.028/1M tokens (cache hit)
European AI gateway that routes to 100+ models with EU data residency
From: Free (10K req/mo), 39 EUR/mo (Plus)
OpenAI-compatible inference, GPU sandboxes and dedicated B200s, hosted in Spa...
From: From EUR 0.06/1M input tokens, EUR 50 free credit on signup
Google's API for Gemini models with text, image, video, and audio capabilities
From: Free
French inference API for open-weight models, hosted on Scaleway with embeddin...
From: 14-day free trial, then EUR 4.50/month
LPU-powered inference API for LLMs, speech, and vision models with usage-base...
From: $0.05/1M tokens
European sovereign AI inference with OpenAI-compatible APIs hosted in EU data...
From: From €0.20/M input tokens (gemma-4-31B)
High-throughput inference API with OpenAI-compatible access to open-source mo...
From: $0.02/M tokens
Search APIs for embeddings, reranking, and web-to-markdown conversion
From: Free (10M tokens)
Build, share and reuse precomputed KV-cache artifacts to skip redundant prefill
EU inference provider serving Qwen3.8-27B from dedicated Helsinki GPUs it ren...
From: $0.25/1M input tokens
Multi-LLM API orchestration platform for comparing and blending AI models
From: $3/300 credits
GPU cloud for AI training and inference with on-demand and cluster options
From: $0.58/GPU/hr (V100)
Web-grounded AI answers API with citations, OpenAI-compatible, pay-per-query ...
From: Free tier, 100 queries/month, then from $3.60/1,000 queries
Run generative AI models, large-scale batch jobs, job queues, and much more.
From: $30/mo free credits
European cloud provider with AI inference, training, and deployment services
From: $0.91/hr (L4 GPU)
OctoAI delivers production-grade GenAI solutions running on the most efficien...
Unified API for 400+ AI models across 60+ providers, OpenAI SDK-compatible, p...
From: Free (25+ free models)
Payment router for AI agents: pay per request across LLM APIs and tools from ...
From: Free tier (BlockRun), then usage-based from $0.001/request
European AI gateway: 700+ models through one EU-hosted, OpenAI-compatible API
From: Provider token rates, no markup; 3% fee on credit purchases
Proves a cheaper model matches your current one on your own prompts, then rou...
From: Free to prove, billed only after activation
Perplexity's open decision model: typed yes/no, choice and score answers with...
From: $0.04 per 1M input tokens
CPU-only LLM inference engine in C with no runtime dependencies
From: Free (open source)
OpenAI-compatible inference API run on Italian infrastructure with zero data ...
LLM gateway and router with one OpenAI-compatible API across 400+ models
From: Free tier, then pay-as-you-go at 5% markup
Unified API for image, video, audio and 3D generation running on custom infer...
From: Pay-as-you-go, $2 free credit (business email)
Custom AI chip inference platform with purpose-built hardware for high-throug...
From: $5 free credit
European serverless AI inference APIs, 100% hosted in Europe
From: From €0.15/M input tokens (chat models)
OpenAI-compatible API serving 200+ open-source LLM and multimodal models
From: Pay-as-you-go, $1 free credit
OpenAI-compatible API for open-weight models, hosted only in EU data centres
OpenAI-compatible inference API for open-weight models, run on EU GPUs with n...
From: $0.12 / 1M input tokens
Unified OpenAI-compatible API gateway to 100+ models across providers
From: From $0.0880/1M tokens (Qwen3.5-plus)
Unified OpenAI-compatible API to 200+ models with smart routing and failover
OpenAI-compatible proxy that trims LLM token usage before requests reach your...
From: From EUR 19/month (10M tokens)
Unified API for hundreds of AI models, with built-in rate limiting and key ma...
From: Free tier, then pay-as-you-go at provider rates
Embedding and reranker models for RAG retrieval quality, from MongoDB
From: Free tier (200M tokens), then from $0.02/1M input tokens
Build the next generation of creativity with fal. Lightning fast inference.
From: $0.02/megapixel
High-throughput LLM inference engine with PagedAttention for efficient GPU me...
From: Free (open-source)
Frequently asked questions
What is the cheapest AI inference API?
On gpt-oss-120b, the open model the most providers here publish a price for, the cheapest per-token rates on the providers' own pricing pages are CoreWeave at $0.03 input and $0.17 output; DeepInfra at $0.037 input and $0.17 output; Melious AI at €0.04 input and €0.20 output (per 1M tokens, checked 25 September 2026). 15 providers are compared in the table on this page. Per-token pricing shifts month to month and varies a lot by model, so check the provider's own pricing page before committing to high-volume workloads.
What is the fastest AI inference API?
It depends on whether sustained throughput or first-token latency matters more. Cerebras reports around 3000 tokens/sec on gpt-oss-120B using WSE hardware, the highest measured throughput in the category as of April 2026. Groq uses custom LPU hardware and runs the same model at ~476 tokens/sec on Artificial Analysis, with a consistently low time-to-first-token (0.6-0.9s) that matters for interactive chat. Both trade off a narrower model catalog than GPU-based providers.
Which AI inference APIs offer a free tier?
Cerebras and Groq both offer free usage with daily token limits, useful for prototyping. Most of the serverless providers (DeepInfra, Together, Fireworks, Novita) hand out free credits on signup rather than a permanent free tier. The "Free tier" filter above lists every provider with a free option.
Which inference providers are OpenAI-compatible?
DeepInfra, Together.ai, Fireworks, Novita, OpenRouter, and Groq all expose a drop-in OpenAI-compatible endpoint. Switching between them usually means changing the base URL and API key, nothing more. Replicate uses its own API format, and raw GPU providers like RunPod and Modal are not endpoints at all, they host whatever gets deployed to them.
Are there EU-hosted AI inference APIs?
Yes. EU-headquartered, GDPR-compliant inference providers include Scaleway (France), Berget AI (Sweden), Cortecs AI (Austria), Infercom (Luxembourg), Tensorix (Ireland), EUrouter (Netherlands), and Lyceum (Germany), all serving inference from European data centres. Use the "European" filter above to see the full list, or visit the European providers page for hosting region details.
What is the best alternative to the OpenAI API?
For the highest sustained throughput on open-source models, Cerebras. For the lowest first-token latency, Groq. For low per-token prices, DeepInfra, CoreWeave and Novita AI list some of the lowest verified rates on gpt-oss-120b (September 2026). For fine-tuning on the same platform as inference, Together.ai or Fireworks. For routing across providers from a single API, OpenRouter.
How to choose an inference API provider
The right provider depends on workload type, latency requirements, and budget. Most providers use pay-per-token pricing for LLMs and per-second GPU billing for custom models. Token-based pricing varies by model, so the cheapest provider for one model may not be cheapest for another.
Free tiers are useful for prototyping but often come with rate limits. For production, compare per-token costs for your specific model, cold start latency, rate limits, and whether the provider supports the models you need.
Teams with data residency requirements should check hosting options and provider headquarters. European providers like Lyceum, 2kw.ai, AKI.IO keep data within EU jurisdiction. See the full European AI Infrastructure directory. Self-hostable options like AISIX and ARK Labs give full control over data location.
For a deeper analysis, read AI Inference API Providers Compared on the blog. Pricing changes frequently, so verify current rates on each provider's website. Submit a correction.
See how these tools fit into a full stack
Browse all Inference APIs tools or explore the full AI Infrastructure Landscape.
Is your product missing?