Airs Bench
AIRS-Bench: an AI Research Science benchmark for quantifying the end-to-end AI research abilities of LLM agents
About
The AI Research Science Benchmark is an eval that quantifies the autonomous research abilities of LLM agents in the area of machine learning. AIRS-Bench comprises 20 tasks from state-of-the-art machine learning papers spanning diverse domains such as NLP, Code, Math, biochemical modelling and time series forecasting.
Open Source Health
- Stars
- 115
- Forks
- 9
- License
- Not stated
- Last commit
- 5 months ago
Resources & Links
Related Categories
Vendor
Facebookresearch
Publisher of Fairseq
Quick Links
Open Source
More by Facebookresearch
Related Products
RPG_KDD2025
This repository provides the code for implementing RPG described in our KDD'25 paper "Generating Long Semantic IDs in Parallel for Recommendation".
More from this vendor
Dagger
the missing software stack for CI
Top category match
Seamless Communication
Foundational Models for State-of-the-Art Speech and Text Translation
More from this vendor
OrienterNet
Source Code for Paper "OrienterNet Visual Localization in 2D Public Maps with Neural Matching"
More from this vendor
Matrix
Matrix (Multi-Agent daTa geneRation Infra and eXperimentation framework) is a versatile engine for multi-agent conversational data generation.
More from this vendor
