Model DeploymentMLOps & Experiment Tracking

Kvpress

LLM KV cache compression made easy

Open Source

About

Deploying long-context LLMs is costly due to the linear growth of the key-value (KV) cache in transformer models. For example, handling 1M tokens with Llama 3.1-70B in float16 requires up to 330GB of memory. kvpress implements multiple KV cache compression methods and benchmarks using 🤗 transformers, aiming to simplify the development of new methods for researchers and developers in this field.

Open Source Health

Not enough history
Stars
1,209
Forks
178
License
Apache-2.0
Last commit
1 months ago
Python

Related Categories