Home / Inference APIs / Compare

Inference APIs Pricing Comparison

118 providers compared by pricing model, free tiers, hosting options, and headquarters. Last updated October 2026.

63 with free tiers · 14 open source · 22 self-hostable · 40 European

AI inference API providers are hosted services that run large language models behind an API, so you pay per token or per second instead of renting and managing GPUs yourself. The table below compares 118 of them on the axes that usually decide the choice: price per token, sustained throughput and first-token latency, the model catalog, OpenAI-compatibility, free tiers, and hosting region (40 run inside the EU). Cost-sensitive and high-throughput workloads tend to pull toward different providers, so the right pick depends on which of these matters most for your use case.

Cheapest gpt-oss-120b providers, ranked by verified price

Per-token price for one model, gpt-oss-120b, across the 15 providers that publish it, so the figures compare like for like. Sorted cheapest first by blended cost (3:1 input to output). The cheapest provider differs by model.

# Provider Input / 1M Cached / 1M Output / 1M Checked Notes
1 $0.03 – $0.17 27 Sep 2026
2 $0.037 – $0.17 25 Sep 2026
3 €0.04 ≈ $0.0455 – €0.20 ≈ $0.2273 1 Oct 2026
4 $0.05 – $0.25 25 Sep 2026
5 $0.05 – $0.45 25 Sep 2026
6 $0.10 – $0.50 25 Sep 2026
7 $0.15 – $0.60 25 Sep 2026
8 $0.15 – $0.60 25 Sep 2026
9 €0.15 ≈ $0.1705 – €0.60 ≈ $0.682 25 Sep 2026
10 €0.15 ≈ $0.1705 – €0.65 ≈ $0.7389 25 Sep 2026 Price on ionos.de; the US site lists $0.17 / $0.71
11 €0.16 ≈ $0.1819 – €0.63 ≈ $0.7161 1 Oct 2026
12 €0.16 ≈ $0.1819 – €0.83 ≈ $0.9435 1 Oct 2026 Listed in SimpleCredits (100 SC = EUR 1); homepage advertises 20% off during the open beta
13 $0.35 – $0.75 25 Sep 2026
14 $0.35 – $0.75 27 Sep 2026 @cf/openai/gpt-oss-120b
15 €1.00 ≈ $1.1367 – €4.20 ≈ $4.7741 1 Oct 2026

Each price was read off the provider's own pricing page on the date shown; the date links to that page. Providers that publish no per-token price for this model are not listed. EUR prices are ranked at the ECB reference rate of 24 September 2026. The same data is open under CC-BY at /api/query/prices.

Prices for DeepSeek, GLM, Kimi, Llama and other models are on the LLM API pricing by model page.

Provider Pricing Model Starting Price Free Tier Hosting Open Source HQ
Pay-per-use From $0.06/1M tokens — Cloud — 🇩🇪 Germany
Pay-per-use Pay per prompt, from $0.10 ✓ Cloud — —
Subscription From EUR 99/month ✓ Cloud — 🇩🇪 Germany
Pay-per-use Free tier, then $0.10/1,000 requests ✓ Cloud — 🇺🇸 United States
Freemium Free ✓ Cloud + Self-hosted ✓ —
Pay-per-use Free (10 EUR credits) ✓ — — 🇩🇪 Germany
Freemium — ✓ Cloud + Self-hosted — —
Subscription 49 SEK/month ✓ Cloud — 🇸🇪 Sweden
Pay-per-use $1.24/hour — Cloud — 🇸🇪 Sweden
Pay-per-use Pay-as-you-go, EUR 10 minimum top-up — Cloud — 🇳🇱 Netherlands
— Free tier (Singapore only), then $2/1M input tokens (Qwen3.8-Max) ✓ — — —
Pay-per-use Pay-per-token ✓ Cloud — 🇺🇸 United States
Pay-per-use $1/1M tokens ✓ Cloud — 🇺🇸 United States
— — — — — 🇺🇸 United States
— From $0.20/1M input tokens — Cloud — 🇵🇱 Poland
Pay-per-use ~$0.63/hr (T4 GPU) ✓ Cloud + Self-hosted — 🇺🇸 United States
Pay-per-use $0.15/hr (T4 GPU) ✓ Cloud + Self-hosted ✓ 🇺🇸 United States
— — ✓ Cloud + Self-hosted ✓ 🇺🇸 United States
Freemium €25/mo ✓ Cloud — 🇸🇪 Sweden
— $0.04/1M input tokens — Cloud — 🇺🇸 United States
Freemium Free tier available ✓ Cloud — 🇺🇸 United States
Pay-per-use ~$1.10/hr (A10 GPU) ✓ Cloud — 🇺🇸 United States
Subscription $22/month (Core pool, one daily 8-hour window) — Cloud — —
Pay-per-use $0.09 per 1M input tokens (Clef-flash on Workers AI) — Cloud + Self-hosted ✓ 🇺🇸 United States
Freemium $0.011/1K neurons ✓ Cloud — 🇺🇸 United States
— — ✓ Cloud — —
Freemium $0.04/1M tokens ✓ Cloud + Self-hosted — 🇨🇦 Canada
Pay-per-use From $0.15 per 1M input tokens — Cloud — 🇺🇸 United States
Pay-per-use $6.50/hr (GH200 GPU) — Cloud — 🇺🇸 United States
Pay-per-use Pay-per-use + 5% gateway fee ✓ Cloud — 🇦🇹 Austria
Pay-per-use $0.02/M tokens ✓ Cloud — 🇺🇸 United States
Pay-per-use $0.028/1M tokens (cache hit) ✓ Cloud ✓ 🇨🇳 China
Subscription Free (10K req/mo), 39 EUR/mo (Plus) ✓ Cloud — 🇳🇱 Netherlands
— — — — — —
— — — — — —
Pay-per-use $0.10/1M tokens — Cloud + Self-hosted — 🇺🇸 United States
— From EUR 0.06/1M input tokens, EUR 50 free credit on signup ✓ — — 🇸🇵🇦🇮🇳 SPAIN
Pay-per-use Pay-as-you-go, $100 free credit on signup — Cloud — 🇺🇸 United States
— — — Cloud — 🇺🇸 United States
Freemium Free ✓ Cloud — 🇺🇸 United States
Subscription 14-day free trial, then EUR 4.50/month ✓ Cloud — 🇳🇱 Netherlands
Freemium $0.05/1M tokens ✓ Cloud — 🇺🇸 United States
Pay-per-use From $0.04 per 1M input tokens — Cloud — 🇸🇰 Slovakia
Pay-per-use From EUR 0.08/1M input tokens, prepaid balance — Cloud + Self-hosted — 🇳🇪🇹🇭🇪🇷🇱🇦🇳🇩🇸 NETHERLANDS
Pay-per-use — — Cloud — 🇦🇪 AE
Pay-per-use $0.15/hr — Cloud — 🇬🇧 United Kingdom
Pay-per-use From $0.11/1M tokens (Qwen3.5-9B) — Cloud — 🇩🇪 Germany
Pay-per-use Prepaid credits, USD 20 minimum funding — Cloud — 🇺🇸 United States
Pay-per-use From €0.20/M input tokens (gemma-4-31B) ✓ Cloud + Self-hosted — 🇱🇺 Luxembourg
Pay-per-use $0.02/M tokens ✓ Cloud — 🇺🇸 United States
Pay-per-use Free (10M tokens) ✓ Cloud — 🇩🇪 Germany
— — ✓ — ✓ —
Kev
Free Free (open-source) — Self-hosted ✓ —
Pay-per-use — — Cloud — 🇩🇪 Germany
Pay-per-use $0.25/1M input tokens ✓ Cloud — 🇵🇱 Poland
— — — — — —
Freemium $3/300 credits ✓ Cloud — 🇺🇸 United States
Pay-per-use $0.58/GPU/hr (V100) ✓ Cloud — 🇺🇸 United States
Free Free (open-source) — Self-hosted ✓ —
Freemium Free tier, then €8/month — Cloud — 🇫🇷 France
Pay-per-use H100 from $1.30/GPU-hour (marketplace rates, Sep 2026) — Cloud — 🇰🇳 KN
Pay-per-use From EUR 49/month — Cloud — 🇩🇪 Germany
Pay-per-use — — Cloud — 🇸🇬 Singapore
Pay-per-use $1.25 / $4.25 per 1M tokens (Muse Spark) — — — 🇺🇸 United States
Pay-per-use Free tier, 100 queries/month, then from $3.60/1,000 queries ✓ Cloud — 🇬🇧 United Kingdom
— — — Cloud — —
Freemium $0.10/1M tokens — Cloud + Self-hosted ✓ 🇫🇷 France
Pay-per-use $30/mo free credits ✓ Cloud — 🇺🇸 United States
Usage Based — — Cloud — —
Pay-per-use $2.00/hr (H100) ✓ Cloud — 🇳🇱 Netherlands
Pay-per-use $0.03/M tokens ✓ Cloud — 🇺🇸 United States
Pay-per-use $0.01/M tokens ✓ Cloud — 🇬🇧 United Kingdom
Pay-per-use $0.91/hr (L4 GPU) ✓ Cloud — 🇫🇷 France
— — ✓ — — 🇺🇸 United States
Pay-per-use $0.05/1M tokens ✓ Cloud — 🇺🇸 United States
Freemium Free (25+ free models) ✓ Cloud — 🇺🇸 United States
Pay-per-use Free tier (BlockRun), then usage-based from $0.001/request ✓ Cloud — —
Pay-per-use Provider token rates, no markup; 3% fee on credit purchases ✓ — — 🇸🇪 Sweden
— — — — — —
Pay-per-use $0.39/hr RTX 4090 (Dedicated) — Cloud — 🇺🇸 United States
Pay-per-use Free to prove, billed only after activation ✓ Cloud + Self-hosted — 🇬🇧 United Kingdom
— $0.004 per 1M tokens (pplx-embed-v1-0.6b) — Cloud — —
— $0.04 per 1M input tokens ✓ Cloud + Self-hosted ✓ —
Contact Sales — — Cloud + Self-hosted — 🇨🇭 Switzerland
Free Free (open source) ✓ Self-hosted ✓ —
— — ✓ — — 🇮🇹 Italy
Pay-per-use Per-second GPU billing — Cloud — 🇺🇸 United States
Freemium Free tier, then pay-as-you-go at 5% markup ✓ Cloud — —
Pay-per-use $0.06/hr — Cloud — 🇺🇸 United States
Pay-per-use Pay-as-you-go, $2 free credit (business email) ✓ Cloud — 🇺🇸 United States
Free Free (open source) — Self-hosted ✓ —
Freemium $5 free credit ✓ Cloud + Self-hosted — 🇺🇸 United States
Pay-per-use From €0.15/M input tokens (chat models) ✓ Cloud — 🇫🇷 France
Pay-per-use Pay-as-you-go, $1 free credit ✓ Cloud — —
— — ✓ — — 🇩🇪 Germany
Subscription From EUR 15/month — Cloud — 🇮🇹 Italy
Pay-per-use $0.12 / 1M input tokens ✓ Cloud — 🇫🇷 France
Pay-per-use $0.0015/image — Cloud — 🇺🇸 United States
Pay-per-use $0.025 per 1M input tokens, output free — — — 🇩🇪 Germany
Pay-per-use ~€2.70/GPU-hr — Cloud — 🇩🇪 Germany
Pay-per-use From $0.02/M tokens — Cloud — 🇮🇪 Ireland
— $0.15/hour — — ✓ —
Pay-per-use Pay-per-token — Cloud + Self-hosted — 🇺🇸 United States
Pay-per-use From $0.0880/1M tokens (Qwen3.5-plus) ✓ — — 🇭🇰 HK
— — ✓ — — —
Pay-per-use $0.042 per 1M input tokens — Cloud — 🇺🇸 United States
— From EUR 19/month (10M tokens) ✓ — — 🇱🇹 Lithuania
Pay-per-use ~$0.06/GPU/hr — Cloud — 🇺🇸 United States
Freemium Free tier, then pay-as-you-go at provider rates ✓ Cloud — 🇺🇸 United States
Pay-per-use $0.14/hr — Cloud — 🇫🇮 Finland
Pay-per-use Free tier (200M tokens), then from $0.02/1M input tokens ✓ Cloud — 🇺🇸 United States
Pay-per-use From $5/month — Cloud — 🇮🇳 India
Subscription From EUR 11.99/month — Cloud — 🇳🇴 Norway
Usage Based — — Cloud — —
Contact Sales — — Cloud — 🇸🇪 Sweden
fal
Pay-per-use $0.02/megapixel ✓ Cloud — 🇺🇸 United States
Free Free (open-source) ✓ Self-hosted ✓ 🇺🇸 United States
— — — Self-hosted — —
ℹ️ Pricing units vary by provider type: per-token for LLM APIs, per-GPU-hour for compute platforms, per-request for media generation. Verify current rates on each provider's website.

Ordering is alphabetical, apart from featured listings, which are paid placement and are labelled as such. Nothing here is ranked by quality. Want to be featured?

Providers with free tiers

These inference apis providers offer free credits, free tiers, or open-source self-hosting options to get started without upfront costs.

One OpenAI-compatible API for 600+ models, with text billed at provider list ...

From: Pay per prompt, from $0.10

EU-hosted AI integration platform combining a provider gateway, document extr...

From: From EUR 99/month

LLM gateway built around spend limits enforced before a request reaches a pro...

From: Free tier, then $0.10/1,000 requests

Open-source AI gateway in Rust with one OpenAI-compatible API across 100+ LLM...

From: Free

European AI API for open-source models on EU infrastructure

From: Free (10 EUR credits)

Sovereign AI inference infrastructure for regulated EU environments, with het...

Show all 63 providers with free tiers

Swedish GPU infrastructure and LLM hosting platform with API-first deployment...

From: 49 SEK/month

Hosted API access to Qwen and third-party models across six global regions

From: Free tier (Singapore only), then $2/1M input tokens (Qwen3.8-Max)

Managed API access to foundation models on AWS with built-in fine-tuning and ...

From: Pay-per-token

Claude API for building AI applications with Opus, Sonnet, and Haiku models

From: $1/1M tokens

AI inference platform for deploying and serving ML models with autoscaling an...

From: ~$0.63/hr (T4 GPU)

Open-source serverless GPU cloud with sub-second cold starts and auto-scaling

From: $0.15/hr (T4 GPU)

BentoML is the platform for software engineers to build AI products.

EU-sovereign AI inference platform with OpenAI-compatible API

From: €25/mo

Ultra-fast inference on custom wafer-scale hardware with OpenAI-compatible API

From: Free tier available

Serverless GPU infrastructure for deploying AI models with sub-5 second cold ...

From: ~$1.10/hr (A10 GPU)

Run AI models at the edge on Cloudflare's global network with serverless infe...

From: $0.011/1K neurons

Unified AI API gateway providing access to 600+ models from OpenAI, Anthropic...

Cohere’s world-class LLMs help enterprises build powerful, secure application...

From: $0.04/1M tokens

European AI inference gateway with smart routing across EU providers

From: Pay-per-use + 5% gateway fee

Run the top AI models using a simple API, pay per use. Low cost, scalable and...

From: $0.02/M tokens

Cost-effective inference API with OpenAI-compatible endpoints and open-weight...

From: $0.028/1M tokens (cache hit)

European AI gateway that routes to 100+ models with EU data residency

From: Free (10K req/mo), 39 EUR/mo (Plus)

OpenAI-compatible inference, GPU sandboxes and dedicated B200s, hosted in Spa...

From: From EUR 0.06/1M input tokens, EUR 50 free credit on signup

Google's API for Gemini models with text, image, video, and audio capabilities

From: Free

French inference API for open-weight models, hosted on Scaleway with embeddin...

From: 14-day free trial, then EUR 4.50/month

LPU-powered inference API for LLMs, speech, and vision models with usage-base...

From: $0.05/1M tokens

European sovereign AI inference with OpenAI-compatible APIs hosted in EU data...

From: From €0.20/M input tokens (gemma-4-31B)

High-throughput inference API with OpenAI-compatible access to open-source mo...

From: $0.02/M tokens

Search APIs for embeddings, reranking, and web-to-markdown conversion

From: Free (10M tokens)

Build, share and reuse precomputed KV-cache artifacts to skip redundant prefill

EU inference provider serving Qwen3.8-27B from dedicated Helsinki GPUs it ren...

From: $0.25/1M input tokens

Multi-LLM API orchestration platform for comparing and blending AI models

From: $3/300 credits

GPU cloud for AI training and inference with on-demand and cluster options

From: $0.58/GPU/hr (V100)

Web-grounded AI answers API with citations, OpenAI-compatible, pay-per-query ...

From: Free tier, 100 queries/month, then from $3.60/1,000 queries

Run generative AI models, large-scale batch jobs, job queues, and much more.

From: $30/mo free credits

Full-stack AI cloud with GPU infrastructure for training and inference

From: $2.00/hr (H100)

APIs, Serverless and GPU Instance In One AI Cloud

From: $0.03/M tokens

European AI hyperscaler with serverless inference and GPU cloud

From: $0.01/M tokens

European cloud provider with AI inference, training, and deployment services

From: $0.91/hr (L4 GPU)

OctoAI delivers production-grade GenAI solutions running on the most efficien...

API access to GPT, o-series reasoning, DALL-E, and Whisper models

From: $0.05/1M tokens

Unified API for 400+ AI models across 60+ providers, OpenAI SDK-compatible, p...

From: Free (25+ free models)

Payment router for AI agents: pay per request across LLM APIs and tools from ...

From: Free tier (BlockRun), then usage-based from $0.001/request

European AI gateway: 700+ models through one EU-hosted, OpenAI-compatible API

From: Provider token rates, no markup; 3% fee on credit purchases

Proves a cheaper model matches your current one on your own prompts, then rou...

From: Free to prove, billed only after activation

Perplexity's open decision model: typed yes/no, choice and score answers with...

From: $0.04 per 1M input tokens

CPU-only LLM inference engine in C with no runtime dependencies

From: Free (open source)

OpenAI-compatible inference API run on Italian infrastructure with zero data ...

LLM gateway and router with one OpenAI-compatible API across 400+ models

From: Free tier, then pay-as-you-go at 5% markup

Unified API for image, video, audio and 3D generation running on custom infer...

From: Pay-as-you-go, $2 free credit (business email)

Custom AI chip inference platform with purpose-built hardware for high-throug...

From: $5 free credit

European serverless AI inference APIs, 100% hosted in Europe

From: From €0.15/M input tokens (chat models)

OpenAI-compatible API serving 200+ open-source LLM and multimodal models

From: Pay-as-you-go, $1 free credit

OpenAI-compatible API for open-weight models, hosted only in EU data centres

OpenAI-compatible inference API for open-weight models, run on EU GPUs with n...

From: $0.12 / 1M input tokens

Unified OpenAI-compatible API gateway to 100+ models across providers

From: From $0.0880/1M tokens (Qwen3.5-plus)

Unified OpenAI-compatible API to 200+ models with smart routing and failover

OpenAI-compatible proxy that trims LLM token usage before requests reach your...

From: From EUR 19/month (10M tokens)

Unified API for hundreds of AI models, with built-in rate limiting and key ma...

From: Free tier, then pay-as-you-go at provider rates

Embedding and reranker models for RAG retrieval quality, from MongoDB

From: Free tier (200M tokens), then from $0.02/1M input tokens

fal

Build the next generation of creativity with fal. Lightning fast inference.

From: $0.02/megapixel

High-throughput LLM inference engine with PagedAttention for efficient GPU me...

From: Free (open-source)

Frequently asked questions

What is the cheapest AI inference API?

On gpt-oss-120b, the open model the most providers here publish a price for, the cheapest per-token rates on the providers' own pricing pages are CoreWeave at $0.03 input and $0.17 output; DeepInfra at $0.037 input and $0.17 output; Melious AI at €0.04 input and €0.20 output (per 1M tokens, checked 25 September 2026). 15 providers are compared in the table on this page. Per-token pricing shifts month to month and varies a lot by model, so check the provider's own pricing page before committing to high-volume workloads.

What is the fastest AI inference API?

It depends on whether sustained throughput or first-token latency matters more. Cerebras reports around 3000 tokens/sec on gpt-oss-120B using WSE hardware, the highest measured throughput in the category as of April 2026. Groq uses custom LPU hardware and runs the same model at ~476 tokens/sec on Artificial Analysis, with a consistently low time-to-first-token (0.6-0.9s) that matters for interactive chat. Both trade off a narrower model catalog than GPU-based providers.

Which AI inference APIs offer a free tier?

Cerebras and Groq both offer free usage with daily token limits, useful for prototyping. Most of the serverless providers (DeepInfra, Together, Fireworks, Novita) hand out free credits on signup rather than a permanent free tier. The "Free tier" filter above lists every provider with a free option.

Which inference providers are OpenAI-compatible?

DeepInfra, Together.ai, Fireworks, Novita, OpenRouter, and Groq all expose a drop-in OpenAI-compatible endpoint. Switching between them usually means changing the base URL and API key, nothing more. Replicate uses its own API format, and raw GPU providers like RunPod and Modal are not endpoints at all, they host whatever gets deployed to them.

Are there EU-hosted AI inference APIs?

Yes. EU-headquartered, GDPR-compliant inference providers include Scaleway (France), Berget AI (Sweden), Cortecs AI (Austria), Infercom (Luxembourg), Tensorix (Ireland), EUrouter (Netherlands), and Lyceum (Germany), all serving inference from European data centres. Use the "European" filter above to see the full list, or visit the European providers page for hosting region details.

What is the best alternative to the OpenAI API?

For the highest sustained throughput on open-source models, Cerebras. For the lowest first-token latency, Groq. For low per-token prices, DeepInfra, CoreWeave and Novita AI list some of the lowest verified rates on gpt-oss-120b (September 2026). For fine-tuning on the same platform as inference, Together.ai or Fireworks. For routing across providers from a single API, OpenRouter.

How to choose an inference API provider

The right provider depends on workload type, latency requirements, and budget. Most providers use pay-per-token pricing for LLMs and per-second GPU billing for custom models. Token-based pricing varies by model, so the cheapest provider for one model may not be cheapest for another.

Free tiers are useful for prototyping but often come with rate limits. For production, compare per-token costs for your specific model, cold start latency, rate limits, and whether the provider supports the models you need.

Teams with data residency requirements should check hosting options and provider headquarters. European providers like Lyceum, 2kw.ai, AKI.IO keep data within EU jurisdiction. See the full European AI Infrastructure directory. Self-hostable options like AISIX and ARK Labs give full control over data location.

For a deeper analysis, read AI Inference API Providers Compared on the blog. Pricing changes frequently, so verify current rates on each provider's website. Submit a correction.

Browse all Inference APIs tools or explore the full AI Infrastructure Landscape.

Is your product missing?

Add it here →