Model DeploymentMLOps & Experiment Tracking

SpAtten

[HPCA'21] SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning

Open Source

About

We propose sparse attention (SpAtten) with KV token pruning, local V pruning, head pruning, and KV progressive quantization to improve LLM efficiency.

Open Source Health

Not enough history
Stars
140
Forks
11
License
MIT
Last commit
2 years ago
Scala

Related Categories