Marketplace

By Use Case Marketplace

Compare By Use Case products across curated subcategories and trusted providers

Products
15,460
Subcategories
699

Selected subcategory

Model Deployment

Compare By Use Case products across curated subcategories and trusted providers

Explore By Use Case with structured category paths, practical filters, and independent product data

Showing 25–36 of 76

Active filtersOpen sourceClear all

AlphaGenome

by Google DeepMind

A Python SDK for interacting and visualizing genomic models.

Model DeploymentMLOps & Experiment TrackingOpen Source

Model export recipes, Python primitives, and Swift runtime utilities for on-device AI

Model DeploymentMLOps & Experiment TrackingOpen Source

h3.c

by Salvatore Sanfilippo

MiniMax H3 inference engine for Mac computers

Model DeploymentMLOps & Experiment TrackingOpen Source

xLLM

by xLLM-AI

A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

Model DeploymentMLOps & Experiment TrackingOpen Source

NInfer

by HaoJun ZHANG

High-performance single-GPU inference for selected model checkpoints and GPUs.

Model DeploymentMLOps & Experiment TrackingOpen Source

GPT4All

by Nomic AI

GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.

Model DeploymentMLOps & Experiment TrackingOpen Source

Solar

by Nvlabs

Speed of Light Analysis for ML Model Runtime

Model DeploymentMLOps & Experiment TrackingOpen Source

NeMo-Speech.cpp is a lightweight C++ inference runtime for Speech models

Model DeploymentMLOps & Experiment TrackingOpen Source

LWS

by Kubernetes Sigs

LeaderWorkerSet: An API for deploying a group of pods as a unit of replication

Model DeploymentMLOps & Experiment TrackingOpen Source

TorchSparse

by MIT HAN Lab

[MICRO'23, MLSys'22] TorchSparse: Efficient Training and Inference Framework for Sparse Convolution on GPUs.

Model DeploymentMLOps & Experiment TrackingOpen Source

SpAtten

by MIT HAN Lab

[HPCA'21] SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning

Model DeploymentMLOps & Experiment TrackingOpen Source

OmniServe

by MIT HAN Lab

[MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention

Model DeploymentMLOps & Experiment TrackingOpen Source

Frequently asked questions