AI Agent PlatformOpen Source AI & Machine Learning

LLM Speedrunner

The Automated LLM Speedrunning Benchmark measures how well LLM agents can reproduce previous innovations and discover new ones in language modeling.

Open Source

About

The Automated LLM Speedrunning Benchmark turns the NanoGPT Speedrun into an eval for the ability of frontier LLM agents to reproduce scientific findings. In this benchmark, an LLM agent is tasked with reproducing the innovations behind each of the NanoGPT Speedrun records, when given access to a description of the innovation. These hints can be set to any of three formats: pseudocode of the change (level 1), text description (level 2), markdown paper describing the improvement (level 3). Currently, no frontier model is capable of reproducing the human-driven speedrun, even when given the pseudocode hints.

Open Source Health

Not enough history
Stars
145
Forks
15
License
Not stated
Last commit
5 months ago
Jupyter Notebook

Related Categories