PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai
☆1,316Jul 2, 2026Updated last month
Alternatives and similar repositories for skill
Users that are interested in skill are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.☆747Updated this week
- An in-the-wild benchmark for AI agents in the OpenClaw Environment.☆510Updated this week
- General Agent Benchmark for OpenClaw, made by Qwen Team, Alibaba Group.☆59Jun 10, 2026Updated 2 months ago
- SkillsBench evaluates how well skills work and how effective agents are at using them.☆1,679Jul 23, 2026Updated 3 weeks ago
- OpenClaw-RL: Train any agent simply by talking☆5,633May 23, 2026Updated 2 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution☆450Updated this week
- A benchmark for LLMs on complicated tasks in the terminal☆2,534Jul 11, 2026Updated last month
- Skill + Plugin Registry for OpenClaw☆9,301Updated this week
- Lossless Claw — LCM (Lossless Context Management) plugin for OpenClaw☆4,886Updated this week
- Framework for evaluating and improving agents☆4,186Updated this week
- τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains☆1,791Updated this week
- ☆33Jun 30, 2026Updated last month
- The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞☆51,928Updated this week
- Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞