Packet.ai
On-demand NVIDIA GPU cloud with per-second billing, SSH, CLI, and API access
Packet.ai is an on-demand GPU cloud for AI and ML workloads, built by hosted.ai and headquartered in San Jose, California. It offers NVIDIA B200, A100, RTX 6000 Pro, RTX 4090, and L40S GPUs on Dedicated (single-tenant) or Dynamic (shared, scheduler-isolated) plans, with full root SSH, a CLI, and API access.
Rates start at USD 0.39/hr for an RTX 4090 (Dedicated) up to USD 5.90/hr for a B200 (Dedicated), billed per second with no long-term contracts. Token Factory, an OpenAI-compatible per-token inference API, is announced but not yet live.
Pricing: Pay-as-you-go
Packet.ai Alternatives
Explore 96 products in the Inference APIs category. View all Packet.ai alternatives.
AI Gateway HQ
LLM gateway built around spend limits enforced before a request reaches a provider
Cloudflare Workers AI
Run AI models at the edge on Cloudflare's global network with serverless inference
novita.ai
APIs, Serverless and GPU Instance In One AI Cloud
Replicate
Run and fine-tune open-source models. Deploy custom models at scale. All with one line of code.
fal
Build the next generation of creativity with fal. Lightning fast inference.
Work on Packet.ai? Feature it at the top of Inference APIs.
Is your product missing?