MInference
MInference accelerates the pre-filling stage of long-context LLM inference with dynamic sparse attention, reaching up to 10x speedup for 1M-token prompts on a single A100. NeurIPS 2024 Spotlight.
About
[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.
Open Source Health
- Stars
- 1,228
- Forks
- 82
- License
- MIT
- Last commit
- 28 days ago
Related Categories
Vendor
Microsoft
Empower every person and organization to achieve more.
Quick Links
Open Source
More by Microsoft
Related Products
Recognizers-Text
Recognizers-Text offers entity recognition and resolution for numbers, dates, and units in multiple languages.
More from this vendor
Language Server Protocol
Defines a common protocol for language servers.
More from this vendor
Lamar Benchmark
Source code for the ECCV 2022 paper "Benchmarking Localization and Mapping for Augmented Reality".
More from this vendor
Pyright
Static Type Checker for Python
More from this vendor
Component Detection
Scans your project to determine what components you use
More from this vendor