Evals
LLM OrchestrationOpen Source AI & Machine Learning

Evals

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

Open Source

About

Evals provide a framework for evaluating large language models (LLMs) or systems built using LLMs. We offer an existing registry of evals to test different dimensions of OpenAI models and the ability to write your own custom evals for use cases you care about. You can also use your data to build private evals which represent the common LLMs patterns in your workflow without exposing any of that data publicly.

Open Source Health

Not enough history
Stars
19,444
Forks
3,083
License
Not stated
Last commit
6 months ago
Python

Related Categories