PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai
☆1,334Jul 2, 2026Updated 2 months ago
Alternatives and similar repositories for skill
Users that are interested in skill are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.☆762Updated this week
- An in-the-wild benchmark for AI agents in the production harness.☆516Aug 17, 2026Updated 2 weeks ago
- General Agent Benchmark for OpenClaw, made by Qwen Team, Alibaba Group.☆59Jun 10, 2026Updated 2 months ago
- SkillsBench evaluates how well skills work and how effective agents are at using them.☆1,741Jul 23, 2026Updated last month
- OpenClaw-RL: Train any agent simply by talking☆5,665May 23, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution☆477Aug 18, 2026Updated 2 weeks ago
- A benchmark for LLMs on complicated tasks in the terminal☆2,560Jul 11, 2026Updated last month
- Skill + Plugin Registry for OpenClaw☆9,386Updated this week
- Lossless Claw — LCM (Lossless Context Management) plugin for OpenClaw☆4,895Updated this week
- Framework for evaluating and improving agents☆4,876Updated this week
- τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains☆1,938Updated this week
- ☆42Jun 30, 2026Updated 2 months ago
- The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞☆52,324Updated this week
- Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞☆388,630Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- SWE-bench: Can Language Models Resolve Real-world Github Issues?☆5,765Updated this week
- 🦞 Just talk to your agent — it learns and EVOLVES 🧬.☆3,496Jun 7, 2026Updated 2 months ago
- The agent that grows with you☆240,119Updated this week
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents☆822May 11, 2026Updated 3 months ago
- An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, s…☆81,265Updated this week
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆23,258Updated this week
- [ACL2026 Main] AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts☆94Jan 23, 2026Updated 7 months ago
- "CLI-Anything: Making ALL Software Agent-Native" -- CLI-Hub: https://clianything.cc/☆48,857Aug 21, 2026Updated last week
- A community collection of OpenClaw use cases for making life easier.☆31,673Mar 24, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, …☆47,654Updated this week
- Democratizing Reinforcement Learning for LLMs☆5,813Aug 24, 2026Updated last week
- The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.☆5,712Updated this week
- slime is an LLM post-training framework for RL Scaling.☆8,354Updated this week
- Public repository for Agent Skills☆173,203Updated this week
- AI agents running research on single-GPU nanochat training automatically☆95,118Mar 26, 2026Updated 5 months ago
- Tongyi Deep Research, the Leading Open-source Deep Research Agent☆19,906Feb 27, 2026Updated 6 months ago
- Open-source Environment toolkit of claw-like agents, support task/harness generation and evaluation☆60May 7, 2026Updated 3 months ago
- AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI☆100,922Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents☆22Jun 9, 2026Updated 2 months ago
- A high-throughput and memory-efficient inference and serving engine for LLMs☆90,787Updated this week
- Make Any Website into CLI & Use your logged-in browser by AI agent.☆28,899Updated this week
- Agent framework and applications built upon Qwen>=3.0, featuring Function Calling, MCP, Code Interpreter, RAG, Chrome extension, etc.☆17,055Mar 4, 2026Updated 5 months ago
- Code and Data for Tau-Bench☆1,419Mar 18, 2026Updated 5 months ago
- Secure, Fast, and Extensible Sandbox runtime for AI agents.☆14,914Updated this week
- Self-evolving memory OS for LLM & AI Agents: ultra-persistent memory, hybrid-retrieval, and cross-task skill reuse, with 35.24% token sav…☆11,158Updated this week