TurboQuant
Model DeploymentMLOps & Experiment Tracking

TurboQuant

TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration

Open Source

About

Implementation of TurboQuant KV cache compression (ICLR 2026, arXiv:2504.19874) with vLLM integration. Tested on dense and MoE architectures across RTX 3090 and RTX 5090 GPUs.

Open Source Health

Not enough history
Stars
1,767
Forks
196
License
GPL-3.0
Last commit
1 months ago
Python

Related Categories