Marketplace
By Use Case Marketplace
Compare By Use Case products across curated subcategories and trusted providers
- Products
- 15,413
- Subcategories
- 698
Selected subcategory
AI Evaluation
Compare By Use Case products across curated subcategories and trusted providers
Explore By Use Case with structured category paths, practical filters, and independent product data
Showing 1–10 of 10
by Confident Ai
The LLM Evaluation Framework
by Anton Obukhov
High-fidelity performance metrics for generative models in PyTorch
by OpenGVLab
[ICLR 2026 Oral] ScaleCUA is the open-sourced computer use agents that can operate on cross-platform environments (Windows, macOS, Ubuntu, Android).
by JonathonLuiten
HOTA (and other) evaluation metrics for Multi-Object Tracking (MOT).
by Microsoft
Prompty is an asset class and format for LLM prompts
by Hugging Face
A benchmark for AI-driven CAD generation and editing
by Hugging Face
Build, enrich, and transform datasets using AI models with no code
by Facebookresearch
PArametrized Recommendation and Ai Model benchmark is a repository for development of numerous uBenchmarks as well as end to end nets for evaluation of training and inference platforms.
by Facebookresearch
AssemblyHands Toolkit is a Python package that provides data loader, visualization, and evaluation tools for the AssemblyHands dataset (CVPR 2023).