LLM Speedrunner
The Automated LLM Speedrunning Benchmark measures how well LLM agents can reproduce previous innovations and discover new ones in language modeling.
About
The Automated LLM Speedrunning Benchmark turns the NanoGPT Speedrun into an eval for the ability of frontier LLM agents to reproduce scientific findings. In this benchmark, an LLM agent is tasked with reproducing the innovations behind each of the NanoGPT Speedrun records, when given access to a description of the innovation. These hints can be set to any of three formats: pseudocode of the change (level 1), text description (level 2), markdown paper describing the improvement (level 3). Currently, no frontier model is capable of reproducing the human-driven speedrun, even when given the pseudocode hints.
Open Source Health
- Stars
- 145
- Forks
- 15
- License
- Not stated
- Last commit
- 5 months ago
Resources & Links
Related Categories
Vendor
Facebookresearch
Publisher of Fairseq
Quick Links
Open Source
Stars
145
Forks
15
Contributors
0
Open Issues
10
More by Facebookresearch
Related Products
RPG_KDD2025
This repository provides the code for implementing RPG described in our KDD'25 paper "Generating Long Semantic IDs in Parallel for Recommendation".
More from this vendor
Dagger
the missing software stack for CI
Top category match
Seamless Communication
Foundational Models for State-of-the-Art Speech and Text Translation
More from this vendor
OrienterNet
Source Code for Paper "OrienterNet Visual Localization in 2D Public Maps with Neural Matching"
More from this vendor
Matrix
Matrix (Multi-Agent daTa geneRation Infra and eXperimentation framework) is a versatile engine for multi-agent conversational data generation.
More from this vendor