Sarathi Serve
A low-latency & high-throughput serving engine for LLMs
About
Sarathi-Serve is a high througput and low-latency LLM serving framework. Please refer to our OSDI'24 paper for more details.
Open Source Health
- Stars
- 522
- Forks
- 67
- License
- Apache-2.0
- Last commit
- 9 months ago
Resources & Links
Related Categories
Vendor
Microsoft
Empower every person and organization to achieve more.
Quick Links
Open Source
More by Microsoft
Related Products
BitNet
Official inference framework for 1-bit LLMs
More from this vendor
CHaiDNN
HLS based Deep Neural Network Accelerator Library for Xilinx Ultrascale+ MPSoCs
Top category match
SoulX-FlashTalk
SoulX-FlashTalk is the first 14B model to achieve sub-second start-up latency (0.87s) while maintaining a real-time throughput of 32 FPS on an 8xH800 node.
Top category match
Ts Pattern
The exhaustive Pattern Matching library for TypeScript, with smart type inference.
Top category match
Nobodywho
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
Top category match
