SimpleLLM
OpenAI-compatible API for open-weight models, hosted only in EU data centres
SimpleLLM is a German-run inference API serving open-weight text, vision and image models through an OpenAI-compatible endpoint. Official SDKs cover JavaScript, Rust, Go and C++.
It also handles fine-tuning, with LoRA, QLoRA, SFT and DPO runs where the GPU provisioning is managed for you.
The positioning is jurisdictional: inference runs on EU providers rather than AWS, GCP or Azure, with a free tier for testing before committing. Worth noting it rents capacity rather than owning hardware, so the guarantee is about where processing happens, not who owns the racks.
Pricing: Per token usage
SimpleLLM Alternatives
Explore 88 products in the Inference APIs category. View all SimpleLLM alternatives.
vLLM
High-throughput LLM inference engine with PagedAttention for efficient GPU memory usage
Genesis Cloud
European GPU cloud, website offline and company in liquidation as of August 2026
Infer by Flow7
Responses API gateway for coding agents, with a public model catalog and per-request spend ceilings
TensorX
EU-sovereign inference API with 42+ open-source models and zero data retention
EUrouter
European AI gateway that routes to 100+ models with EU data residency
Work on SimpleLLM? Feature it at the top of Inference APIs.
Is your product missing?