Icon for Boundless

Boundless

Per-token API for open and frontier models, with engineers who tune the serving setup to your traffic

Boundless is an inference provider run by Boundless Networks, Inc. It serves around 20 models through an OpenAI-compatible API: open models from DeepSeek, Z.ai, Moonshot, MiniMax, Qwen and Xiaomi, plus Claude and GPT. It also hosts the Laya decision model on a /v1/systemone endpoint, which is the route Vercel AI Gateway uses for Laya.

Billing is prepaid credits at per-token rates, from $0.04/1M input on DeepSeek V4 Flash, with cached-input prices on most models. Its platform terms say prompts and outputs are not retained. A forward-deployed option has Boundless engineers choose the model and tune serving for a team's workload.

Pricing: Per token usage

Hosting Cloud
Pricing $0.04/1M input tokens
HQ 🇺🇸 United States
Screenshot of Boundless webpage

Boundless prices by model

Per 1M tokens, read off Boundless's own pricing page on the date shown.

Model Input / 1M Cached / 1M Output / 1M Checked Notes
DeepSeek V4 Pro $1.04 $0.0352 $2.08 3 Oct 2026 Listed as the 0813 snapshot of DeepSeek V4 Pro.
DeepSeek V4.1 Flash $0.20 $0.01 $1.00 3 Oct 2026
GLM 5.3 $1.12 $0.14 $3.52 3 Oct 2026
GLM 5.3 Flash $0.12 $0.024 $0.40 3 Oct 2026
Kimi K3 $2.30 $0.23 $11.40 3 Oct 2026
MiniMax M3 $0.28 $0.056 $1.10 3 Oct 2026

Work on Boundless? Feature it at the top of Inference APIs.

Is your product missing?

Add it here →