Shellbench
LLM OrchestrationOpen Source AI & Machine Learning

Shellbench

The agent benchmark that scores the full stack — harness, config, and model — not just the LLM. Trace-based scoring, reliability metrics, configuration diagnostics.

Open Source

About

title: ClawBench emoji: 🦞 colorFrom: red colorTo: yellow sdk: docker app_port: 7860 pinned: true license: mit

Open Source Health

Not enough history
Stars
139
Forks
29
License
MIT
Last commit
1 months ago
Python

Related Categories