BREAKING
LMCache Speeds LLM Inference Up to 14x
0
x
Faster TTFT
0
x
Higher throughput
0
%
Cost reduction
How LMCache Reuses the KV Cache
1
Compute KV cache
↓
2
Offload to CPU/disk
↓
3
Persist and share
↓
4
Reuse on match
Why LMCache Stands Apart
LMCache
Open-source
●
Vendor-neutral
●
Cross-engine sharing
●
vLLM, SGLang, TensorRT-LLM
Vendor caches
Closed
●
NVIDIA Dynamo KVBM
●
Vendor-specific tokens
●
Engine-locked
0
GitHub stars
0
PyPI downloads
0
Contributors
Open KV Cache for Faster Inference
AI NEWS BLITZ
An open-source KV cache layer is speeding up LLM inference by up to fourteen times.