KV Cache Store
Build, share and reuse precomputed KV-cache artifacts to skip redundant prefill
KV Cache Store lets you precompute the KV cache for a long prompt once, then load it before inference instead of paying the prefill cost on every request. The open-source kvcdn CLI builds, verifies, quantizes and benchmarks cache artifacts locally.
An optional hosted registry stores artifacts behind stable URLs with SHA-256 digests and model, dtype and tokenizer metadata, so a cache can be shared across machines without ambiguity about what produced it.
Useful when many requests share a large fixed context, such as a long system prompt or a document being queried repeatedly. Currently in public beta.
Pricing: Monthly subscriptions
KV Cache Store Alternatives
Explore 96 products in the Inference APIs category. View all KV Cache Store alternatives.
AI Gateway HQ
LLM gateway built around spend limits enforced before a request reaches a provider
Cloudflare Workers AI
Run AI models at the edge on Cloudflare's global network with serverless inference
novita.ai
APIs, Serverless and GPU Instance In One AI Cloud
Replicate
Run and fine-tune open-source models. Deploy custom models at scale. All with one line of code.
fal
Build the next generation of creativity with fal. Lightning fast inference.
Work on KV Cache Store? Feature it at the top of Inference APIs.
Is your product missing?