LLM OrchestrationOpen Source AI & Machine Learning

Star-Attention

Efficient LLM Inference over Long Sequences

Open Source

About

Star Attention improves the inference time by up to 11x while preserving 97-100% of accuracy. The method is compatible with most Transformer-based LLMs trained with global attention, operating seamlessly out-of-the-box without additional training/finetuning. Furthermore, Star Attention is orthogonal to other optimization methods, including Flash Attention and KV cache compression techniques, allowing for potential combined enhancements.

Open Source Health

Not enough history
Stars
392
Forks
24
License
Apache-2.0
Last commit
1 years ago
Python

Related Categories