ds4
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
About
DwarfStar aims to be the best way to run a few excellent large language models on consumer hardware (that is, hardware that people can actually own). To reach this goal, we are building a small native inference engine optimized first for DeepSeek V4 Flash (including the experimental vision model), DeepSeek V4.1 Flash (Metal only), and additionally GLM 5.2 and 5.3, GLM 5.3 Flash and DeepSeek V4 PRO. The code is self-contained and deliberately narrow, not a general GGUF runner: you need to use the GGUF files the project produces, that are part of the project itself.
Open Source Health
- Stars
- 22,332
- Forks
- 2,107
- License
- MIT
- Last commit
- 28 days ago
Resources & Links
Related Categories
Vendor
Salvatore Sanfilippo
Computer programmer based in Sicily, Italy. I mostly write OSS software. Born 1977. Not a puritan
Quick Links
Open Source
More by Salvatore Sanfilippo
Related Products
h3.c
MiniMax H3 inference engine for Mac computers
More from this vendor
CHaiDNN
HLS based Deep Neural Network Accelerator Library for Xilinx Ultrascale+ MPSoCs
Top category match
SoulX-FlashTalk
SoulX-FlashTalk is the first 14B model to achieve sub-second start-up latency (0.87s) while maintaining a real-time throughput of 32 FPS on an 8xH800 node.
Top category match
Ts Pattern
The exhaustive Pattern Matching library for TypeScript, with smart type inference.
Top category match
Nobodywho
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
Top category match