Evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
About
Evals provide a framework for evaluating large language models (LLMs) or systems built using LLMs. We offer an existing registry of evals to test different dimensions of OpenAI models and the ability to write your own custom evals for use cases you care about. You can also use your data to build private evals which represent the common LLMs patterns in your workflow without exposing any of that data publicly.
Open Source Health
- Stars
- 19,444
- Forks
- 3,083
- License
- Not stated
- Last commit
- 6 months ago
Resources & Links
Related Categories
Vendor
Openai
Quick Links
Open Source
More by Openai
Related Products
Whisper
Robust Speech Recognition via Large-Scale Weak Supervision
More from this vendor
Skills
Skills Catalog for Codex
More from this vendor
Symphony
Symphony turns project work into isolated, autonomous implementation runs, allowing teams to manage work instead of supervising coding agents.
More from this vendor
Swarm
Educational framework exploring ergonomic, lightweight multi-agent orchestration. Managed by OpenAI Solution team.
More from this vendor
The Algorithm
Source code for the X Recommendation Algorithm
Shared categories
