PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with š¦ by the humans at https://kilo.ai
ā1,350Jul 2, 2026Updated 2 months ago
Alternatives and similar repositories for skill
Users that are interested in skill are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.ā773Updated this week
- An in-the-wild benchmark for AI agents in the production harness.ā523Updated this week
- General Agent Benchmark for OpenClaw, made by Qwen Team, Alibaba Group.ā61Jun 10, 2026Updated 3 months ago
- SkillsBench evaluates how well skills work and how effective agents are at using them.ā1,810Jul 23, 2026Updated last month
- OpenClaw-RL: Train any agent simply by talkingā5,697May 23, 2026Updated 4 months ago
- Proton VPN Special Offer - Get 70% off ⢠AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution