Open Source

LMCache

LMCache is an open-source KV cache engine that helps large language models scale by separating KV cache storage from GPU memory.

Source: GitHub Pricing: Open Source
💻 View Code

About This Project

LMCache reduces Time to First Token (TTFT), improves GPU utilization, eliminates redundant prefill computation, and enables persistent, stateful inference across distributed LLM servers.

Tags

gpu kv-cache llm-inference

Reviews & Ratings

Share your experience

User Reviews (0)

No reviews yet. Be the first to share your experience!