GPU Flow
OpenAI-compatible inference, GPU sandboxes and dedicated B200s, hosted in Spain and billed in euros
GPU Flow is a Spanish AI infrastructure provider running three products on one prepaid balance: Token Factory (an OpenAI and Anthropic SDK-compatible inference API), Sandboxes (SSH dev environments that auto-suspend when idle), and dedicated NVIDIA B200 GPUs by the hour or month.
The model catalogue is open-weight and covers text, reasoning, vision/OCR, speech and music: DeepSeek V4 Pro and Flash, GLM 5.3, Qwen3.8 27B, Gemma 4, Whisper large-v3, and ALIA 40B, Spain's public model trained in Spanish, Catalan, Basque and Galician.
Compute runs on bare-metal B200s in a Spanish data centre with no sub-processors outside the EU. The company holds ISO 27001 and ENS Medium, both published as downloadable certificates, and invoices in euros with intra-EU reverse charge.
Pricing: Pay-as-you-go
GPU Flow prices by model
Per 1M tokens, read off GPU Flow's own pricing page on the date shown.
| Model | Input / 1M | Cached / 1M | Output / 1M | Checked | Notes |
|---|---|---|---|---|---|
| GLM 5.3 | โฌ1.206 | โ | โฌ3.789 | 1 Oct 2026 | |
| Qwen3.8 27B | โฌ0.10 | โ | โฌ0.30 | 1 Oct 2026 |
GPU Flow Alternatives
Explore 122 products in the Inference APIs category. View all GPU Flow alternatives.
Heabsy
EU inference API on its own GPUs in the EEA, with zero data retention and OpenAI- and Anthropic-compatible endpoints.
DeepInfra
Run the top AI models using a simple API, pay per use. Low cost, scalable and production ready infrastructure.
HostYourAI
EU-hosted inference router with OpenAI and Anthropic drop-in compatibility, plus dedicated vLLM instances
Lium
GPU rental marketplace where independent providers list verified NVIDIA hosts, billed per second
Aurora Inference
OpenAI-compatible API for open-weight models, with in-region deployment on reserved capacity
Work on GPU Flow? Feature it at the top of Inference APIs.
Is your product missing?