HostYourAI
EU-hosted inference router with OpenAI and Anthropic drop-in compatibility, plus dedicated vLLM instances
HostYourAI is a Dutch inference platform that runs open-weight models on European GPUs. The shared EU Router exposes one OpenAI-compatible endpoint and an Anthropic Messages drop-in, so an existing client works by changing the base URL.
Beyond the router it deploys any Hugging Face model on dedicated GPUs with vLLM, billed per minute rather than per hour, with scale-to-zero when idle. Model Garden lists the catalogue with live warm/cold status, and BYOK lets you attach your own OpenAI, Anthropic, Google or Mistral key with no platform fee.
EU residency is a hard constraint rather than a default: a machine qualifies only if its data centre sits in the EU or EEA, and an EU Sovereignty Mode restricts processing to sub-processors established in the EU. A DPA, a public sub-processor list and a status page are all published.
Pricing: Pay-as-you-go
HostYourAI prices by model
Per 1M tokens, read off HostYourAI's own pricing page on the date shown.
| Model | Input / 1M | Cached / 1M | Output / 1M | Checked | Notes |
|---|---|---|---|---|---|
| DeepSeek V4 Pro | β¬2.01 | β¬0.51 | β¬4.03 | 1 Oct 2026 | |
| DeepSeek V4.1 Flash | β¬0.58 | β¬0.15 | β¬1.73 | 1 Oct 2026 | |
| GLM 5.3 | β¬1.15 | β¬0.29 | β¬4.03 | 1 Oct 2026 | |
| GLM 5.3 Flash | β¬0.23 | β¬0.06 | β¬0.58 | 1 Oct 2026 | |
| Kimi K3 | β¬3.45 | β¬0.86 | β¬17.25 | 1 Oct 2026 | |
| Llama 3.3 70B Instruct | β¬1.04 | β | β¬1.04 | 1 Oct 2026 | |
| MiniMax M3 | β¬0.46 | β¬0.12 | β¬2.30 | 1 Oct 2026 | |
| Qwen3.8 2.4T A95B | β¬2.88 | β¬0.72 | β¬6.90 | 1 Oct 2026 | |
| Qwen3.8 27B | β¬0.35 | β¬0.12 | β¬2.53 | 1 Oct 2026 | |
| gpt-oss-120b | β¬0.16 | β | β¬0.63 | 1 Oct 2026 |
HostYourAI Alternatives
Explore 122 products in the Inference APIs category. View all HostYourAI alternatives.
Heabsy
EU inference API on its own GPUs in the EEA, with zero data retention and OpenAI- and Anthropic-compatible endpoints.
DeepInfra
Run the top AI models using a simple API, pay per use. Low cost, scalable and production ready infrastructure.
Lium
GPU rental marketplace where independent providers list verified NVIDIA hosts, billed per second
Aurora Inference
OpenAI-compatible API for open-weight models, with in-region deployment on reserved capacity
Work on HostYourAI? Feature it at the top of Inference APIs.
Is your product missing?