KV Cache Store
Build, share and reuse precomputed KV-cache artifacts to skip redundant prefill
KV Cache Store lets you precompute the KV cache for a long prompt once, then load it before inference instead of paying the prefill cost on every request. The open-source kvcdn CLI builds, verifies, quantizes and benchmarks cache artifacts locally.
An optional hosted registry stores artifacts behind stable URLs with SHA-256 digests and model, dtype and tokenizer metadata, so a cache can be shared across machines without ambiguity about what produced it.
Useful when many requests share a large fixed context, such as a long system prompt or a document being queried repeatedly. Currently in public beta.
Pricing: Monthly subscriptions
KV Cache Store Alternatives
Explore 104 products in the Inference APIs category. View all KV Cache Store alternatives.
vLLM
High-throughput LLM inference engine with PagedAttention for efficient GPU memory usage
Project Zero
CPU-only LLM inference engine in C with no runtime dependencies
SGLang
High-performance open-source serving framework for LLMs and multimodal models
GreenPT
French inference API for open-weight models, hosted on Scaleway with embeddings, reranking and speech
Runware
Unified API for image, video, audio and 3D generation running on custom inference hardware
Work on KV Cache Store? Feature it at the top of Inference APIs.
Is your product missing?