Marketplace

By Use Case Marketplace

Compare By Use Case products across curated subcategories and trusted providers

Products
15,413
Subcategories
698

Selected subcategory

AI Evaluation

Compare By Use Case products across curated subcategories and trusted providers

Explore By Use Case with structured category paths, practical filters, and independent product data

Showing 1–10 of 10

Active filtersOpen sourceClear all

Deepeval

by Confident Ai

The LLM Evaluation Framework

AI EvaluationMLOps & Experiment TrackingOpen Source

Torch Fidelity

by Anton Obukhov

High-fidelity performance metrics for generative models in PyTorch

AI EvaluationMLOps & Experiment TrackingOpen Source

ScaleCUA

by OpenGVLab

[ICLR 2026 Oral] ScaleCUA is the open-sourced computer use agents that can operate on cross-platform environments (Windows, macOS, Ubuntu, Android).

AI EvaluationMLOps & Experiment TrackingOpen Source

Garak

by NVIDIA

the LLM vulnerability scanner

AI EvaluationMLOps & Experiment TrackingOpen Source

TrackEval

by JonathonLuiten

HOTA (and other) evaluation metrics for Multi-Object Tracking (MOT).

AI EvaluationMLOps & Experiment TrackingOpen Source

Prompty

by Microsoft

Prompty is an asset class and format for LLM prompts

AI EvaluationMLOps & Experiment TrackingOpen Source

Cadgenbench

by Hugging Face

A benchmark for AI-driven CAD generation and editing

AI EvaluationMLOps & Experiment TrackingOpen Source

Aisheets

by Hugging Face

Build, enrich, and transform datasets using AI models with no code

AI EvaluationMLOps & Experiment TrackingOpen Source

Param

by Facebookresearch

PArametrized Recommendation and Ai Model benchmark is a repository for development of numerous uBenchmarks as well as end to end nets for evaluation of training and inference platforms.

AI EvaluationMLOps & Experiment TrackingOpen Source

Assemblyhands Toolkit

by Facebookresearch

AssemblyHands Toolkit is a Python package that provides data loader, visualization, and evaluation tools for the AssemblyHands dataset (CVPR 2023).

AI EvaluationMLOps & Experiment TrackingOpen Source

Frequently asked questions