NInfer
High-performance single-GPU inference for selected model checkpoints and GPUs.
About
NInfer is a from-scratch C++/CUDA inference engine for explicitly registered Qwen checkpoints on a single NVIDIA GeForce RTX 5090. It runs text, image, and video prompts through a local CLI or OpenAI-/Anthropic-compatible HTTP APIs. The runtime is deliberately specialized: one GPU, one resident model, and a startup-fixed capacity of one to eight active requests.
Open Source Health
- Stars
- 1,613
- Forks
- 302
- License
- Apache-2.0
- Last commit
- 1 months ago
Resources & Links
Related Categories
Vendor
HaoJun ZHANG
C++ / CUDA · High-performance systems · Local inference, graphics, and digital fabrication
Quick Links
Open Source
Related Products
CHaiDNN
HLS based Deep Neural Network Accelerator Library for Xilinx Ultrascale+ MPSoCs
Top category match
SoulX-FlashTalk
SoulX-FlashTalk is the first 14B model to achieve sub-second start-up latency (0.87s) while maintaining a real-time throughput of 32 FPS on an 8xH800 node.
Top category match
Ts Pattern
The exhaustive Pattern Matching library for TypeScript, with smart type inference.
Top category match
Nobodywho
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
Top category match
ds4
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
Top category match