SimpleLLM
OpenAI-compatible API for open-weight models, hosted only in EU data centres
SimpleLLM is a German-run inference API serving open-weight text, vision and image models through an OpenAI-compatible endpoint. Official SDKs cover JavaScript, Rust, Go and C++.
It also handles fine-tuning, with LoRA, QLoRA, SFT and DPO runs where the GPU provisioning is managed for you.
The positioning is jurisdictional: inference runs on EU providers rather than AWS, GCP or Azure, with a free tier for testing before committing. Worth noting it rents capacity rather than owning hardware, so the guarantee is about where processing happens, not who owns the racks.
Pricing: Per token usage
SimpleLLM Alternatives
Explore 96 products in the Inference APIs category. View all SimpleLLM alternatives.
AI Gateway HQ
LLM gateway built around spend limits enforced before a request reaches a provider
Cloudflare Workers AI
Run AI models at the edge on Cloudflare's global network with serverless inference
novita.ai
APIs, Serverless and GPU Instance In One AI Cloud
Replicate
Run and fine-tune open-source models. Deploy custom models at scale. All with one line of code.
fal
Build the next generation of creativity with fal. Lightning fast inference.
Work on SimpleLLM? Feature it at the top of Inference APIs.
Is your product missing?