LLM Tech
EU inference provider serving Qwen3.8-27B from dedicated hardware in Helsinki, with zero data retention
LLM Tech is an EU inference provider serving a single model, Qwen3.8-27B in NVFP4, from its own dedicated hardware rather than reselling capacity. GPUs sit in Helsinki with the edge in Nuremberg, and no hop uses US cloud.
The API is OpenAI-compatible with streaming, tool calling, structured outputs, image input and reasoning control, on a 262,144-token context and 32,768 max output. Rates are $0.25/1M input and $2.09/1M output, with cached input at $0.04.
Prompts and completions are never written to disk, logs or backups, and an Art. 28 GDPR DPA is available to sign. The model catalogue is public without a key, and uptime and latency are published live.
Pricing: Per token usage
What is LLM Tech?
LLM Tech is a single-model inference provider running unsloth/Qwen3.8-27B-NVFP4 on its own NVIDIA RTX PRO 6000 Blackwell hardware. It is not a router or a reseller: one model, one deployment, tuned for it. The service has been in production since 22 August 2026.
Where it runs
GPUs are in Helsinki, Finland; the edge is in Nuremberg, Germany, over an encrypted tunnel, TLS 1.3 only. The operator states no US cloud is involved at any hop and nothing leaves the EU/EEA. Infrastructure is Hetzner for the edge and Verda for the GPUs.
The API
OpenAI-compatible chat completions at https://api.llmtech.eu/v1, with streaming, tool calling, JSON mode, structured outputs, logprobs, reasoning control and image input. Context is 262,144 tokens with 32,768 max output. The model catalogue at /v1/models is public and needs no key, and a public trial key is published on the quickstart page, so the endpoint can be tested before signing up. Working configs are documented for Cline, Kilo, Continue, aider, Open WebUI, Zed and Claude Code.
Pricing
$0.25 per 1M input tokens, $2.09 per 1M output, $0.04 per 1M cached input. Prompt caching is automatic; the provider reports roughly 88% of production input served from cache, which it puts at an effective input price near $0.07/M. Rates are published in the public model endpoint rather than only on the pricing page.
Data handling
Prompts and completions are never written to disk, logs, analytics or backups, and are not used for training. Metadata is retained 13 months as billing evidence. An Article 28 DPA on the Commission's standard contractual clauses is ready to sign, with two named infrastructure sub-processors, neither having access to request content. The operator confirms GDPR compliance and states no SOC 2 or HIPAA.
Who it is for
Teams with EU data-residency requirements that want one well-tuned open-weight model rather than a broad catalogue. The trade-off is scope and scale: this is a sole proprietorship (Artem Burei, JDG, Poland) running one model, so there is no second model to fall back on and no large-vendor support organisation behind it. Uptime and latency are published live at llmtech.eu/status, measured on production traffic.
LLM Tech Alternatives
Explore 94 products in the Inference APIs category. View all LLM Tech alternatives.
Varion
OpenAI-compatible proxy that trims LLM token usage before requests reach your provider
Scaleway
European serverless AI inference APIs, 100% hosted in Europe
Alibaba Cloud Model Studio
Hosted API access to Qwen and third-party models across six global regions
Work on LLM Tech? Feature it at the top of Inference APIs.
Is your product missing?