LLM OrchestrationOpen Source AI & Machine Learning

MInference

MInference accelerates the pre-filling stage of long-context LLM inference with dynamic sparse attention, reaching up to 10x speedup for 1M-token prompts on a single A100. NeurIPS 2024 Spotlight.

Open Source

About

[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.

Open Source Health

Not enough history
Stars
1,228
Forks
82
License
MIT
Last commit
28 days ago
Python

Related Categories