Infer by Flow7
Responses API gateway for coding agents, with a public model catalog and per-request spend ceilings
Infer by Flow7 is an API gateway that routes requests to models from OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot AI and Qwen through a single key and a prepaid balance.
It exposes a Responses-compatible endpoint at /v1/responses rather than chat completions, and documents verified setups for Codex CLI, OpenCode, Pydantic AI, LangChain and the Vercel AI SDK. Four routing options (Low Cost, Balanced, Stable, and Official API to a model developer's first-party endpoint) let you choose between price and route stability per request.
Spend ceilings are checked before a request is dispatched, and each completed call records a receipt with the resolved model, route tier, token counts and price version. The public catalog lists 21 models with 79 published prices, including labels on the routes that cost more than the model developer's own rate.
Funding starts at a USD 20 minimum, with a USD 2,000 cap on account balance.
Pricing: Per token usage
Infer by Flow7 Alternatives
Explore 96 products in the Inference APIs category. View all Infer by Flow7 alternatives.
AI Gateway HQ
LLM gateway built around spend limits enforced before a request reaches a provider
Cloudflare Workers AI
Run AI models at the edge on Cloudflare's global network with serverless inference
novita.ai
APIs, Serverless and GPU Instance In One AI Cloud
Replicate
Run and fine-tune open-source models. Deploy custom models at scale. All with one line of code.
fal
Build the next generation of creativity with fal. Lightning fast inference.
Work on Infer by Flow7? Feature it at the top of Inference APIs.
Is your product missing?