InfiniStore
Model DeploymentMLOps & Experiment Tracking

InfiniStore

KV cache store for distributed LLM inference

Open Source

About

InfiniStore is an open-source high-performance KV store. It's designed to support LLM Inference clusters, whether the cluster is in prefill-decoding disaggregation mode or not. InfiniStore provides high-performance and low-latency KV cache transfer and KV cache reuse among inference nodes in the cluster.

Open Source Health

Not enough history
Stars
437
Forks
44
License
Apache-2.0
Last commit
11 months ago
C++

Related Categories