Shellbench
The agent benchmark that scores the full stack — harness, config, and model — not just the LLM. Trace-based scoring, reliability metrics, configuration diagnostics.
About
title: ClawBench emoji: 🦞 colorFrom: red colorTo: yellow sdk: docker app_port: 7860 pinned: true license: mit
Open Source Health
- Stars
- 139
- Forks
- 29
- License
- MIT
- Last commit
- 1 months ago
Related Categories
Vendor
openclaw
Your personal, open source AI assistant
Quick Links
Open Source
More by openclaw
Related Products
The Algorithm
Source code for the X Recommendation Algorithm
Shared categories
Recommenders
TensorFlow Recommenders is a library for building recommender system models using TensorFlow.
Shared categories
RPG_KDD2025
This repository provides the code for implementing RPG described in our KDD'25 paper "Generating Long Semantic IDs in Parallel for Recommendation".
Shared categories
Pairec
A Go web framework for quickly building recommendation online services based on JSON configuration.
Shared categories
Textrank
TextRank implementation for Python 3.
Shared categories