NInfer
Model DeploymentMLOps & Experiment Tracking

NInfer

High-performance single-GPU inference for selected model checkpoints and GPUs.

Open Source

About

NInfer is a from-scratch C++/CUDA inference engine for explicitly registered Qwen checkpoints on a single NVIDIA GeForce RTX 5090. It runs text, image, and video prompts through a local CLI or OpenAI-/Anthropic-compatible HTTP APIs. The runtime is deliberately specialized: one GPU, one resident model, and a startup-fixed capacity of one to eight active requests.

Open Source Health

Not enough history
Stars
1,613
Forks
302
License
Apache-2.0
Last commit
1 months ago
C++

Related Categories