Airs Bench
AI Agent PlatformOpen Source AI & Machine Learning

Airs Bench

AIRS-Bench: an AI Research Science benchmark for quantifying the end-to-end AI research abilities of LLM agents

Open Source

About

The AI Research Science Benchmark is an eval that quantifies the autonomous research abilities of LLM agents in the area of machine learning. AIRS-Bench comprises 20 tasks from state-of-the-art machine learning papers spanning diverse domains such as NLP, Code, Math, biochemical modelling and time series forecasting.

Open Source Health

Not enough history
Stars
115
Forks
9
License
Not stated
Last commit
5 months ago
Python

Related Categories