Requesty
LLM gateway and router with one OpenAI-compatible API across 400+ models
Requesty is an LLM gateway that routes requests to 400+ models across 30+ providers through a single OpenAI-compatible endpoint. It handles intelligent routing (by cost, latency, or availability), automatic failover, and prompt caching to cut token spend.
It adds per-user and per-team cost tracking, spend limits, and observability dashboards for cost, latency, error rates, and cache hits. Enterprise plans include SSO, RBAC, audit logs, guardrails, and PII detection. SOC 2, GDPR, and HIPAA compliant, with EU data residency available via a separate EU endpoint.
The free tier covers free models and 200 requests/day; paid usage is a 5% markup on base model costs with bring-your-own-keys.
Pricing: Usage-based
2 developers want to try this
Requesty Alternatives
Explore 90 products in the Inference APIs category. View all Requesty alternatives.
vLLM
High-throughput LLM inference engine with PagedAttention for efficient GPU memory usage
Genesis Cloud
European GPU cloud, website offline and company in liquidation as of August 2026
Infer by Flow7
Responses API gateway for coding agents, with a public model catalog and per-request spend ceilings
Work on Requesty? Feature it at the top of Inference APIs.
Is your product missing?