Repository for results and data (coming soon!) for ClawsBench
☆34Aug 28, 2026Updated 2 weeks ago
Alternatives and similar repositories for ClawsBench
Users that are interested in ClawsBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆20Aug 9, 2026Updated last month
- SkillsBench evaluates how well skills work and how effective agents are at using them.☆1,764Jul 23, 2026Updated last month
- Open-source Environment toolkit of claw-like agents, support task/harness generation and evaluation☆61May 7, 2026Updated 4 months ago
- Tutorial on Benders decomposition and acceleration techniques☆18May 23, 2023Updated 3 years ago
- AI agent benchmark hackability scanner — find evaluation vulnerabilities before they undermine your results☆46May 25, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- An implementation of a device tracking technique based on Algorithm 4 (Double-Hash Port Selection) of RFC 6056.☆16Sep 28, 2022Updated 3 years ago
- ☆18Jan 6, 2025Updated last year
- A new ChatPDF☆14Mar 20, 2025Updated last year
- Code for the arxiv paper: Complex Claim Verification with Evidence Retrieved in the Wild☆14Nov 27, 2023Updated 2 years ago
- SWE-Marathon: an ultra long-horizon SWE benchmark☆152Updated this week
- ☆18Updated this week
- MMDeepResearch-Bench (MMDR)☆32Aug 10, 2026Updated last month
- Measuring and evolving with the frontier of agent work☆673Updated this week
- ResNet-50 for TsinghuaDog classification☆10Feb 2, 2021Updated 5 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆14Jan 6, 2025Updated last year
- [EMNLP 2026] Official repository for Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw☆68May 2, 2026Updated 4 months ago
- [USENIX'25] HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns☆15Mar 1, 2025Updated last year
- 新华网和人民网的简单关键词Scrapy爬虫☆12Jun 2, 2022Updated 4 years ago
- [CVPR 2026] Code for Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective Memory☆17Jul 30, 2026Updated last month
- mcp server for tidb☆24Apr 15, 2025Updated last year
- Flutter VR app using React 360 and GitHub Pages.☆19Nov 27, 2021Updated 4 years ago
- 网络嗅探实验☆16Apr 6, 2022Updated 4 years ago
- [EMNLP 2022] This is the code repo for our EMNLP‘22 paper "Dimension Reduction for Efficient Dense Retrieval via Conditional Autoencoder"…☆13Oct 20, 2022Updated 3 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆13Dec 12, 2024Updated last year
- An arbitrage bot is a smart contract connected to an external automation script that controls its operation.☆2,698Updated this week
- 🦞 ClawMark: A Living-World Benchmark for Multi-Day, Multimodal Coworker Agents☆124May 28, 2026Updated 3 months ago
- Python ESPIRiT implementation☆11Feb 24, 2017Updated 9 years ago
- Multimodal Safety Awareness Benchmark for Large Language Models☆15Jun 3, 2025Updated last year
- ☆19Oct 1, 2025Updated 11 months ago
- Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.☆773Updated this week
- Χ-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?☆61Sep 2, 2026Updated last week
- Codebase for the paper "Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep Learning"☆17Jul 12, 2021Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆24Jul 20, 2026Updated last month
- Official PyTorch Implementation of KL-Tracing: Taming generative video models for zero-shot optical flow extraction.☆19Jul 15, 2025Updated last year
- This repo accompanies the paper "The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context M…☆18Nov 18, 2025Updated 9 months ago
- Extract your SlidesLive presentation.☆15Apr 19, 2024Updated 2 years ago
- A high-throughput oblivious storage system☆29May 31, 2023Updated 3 years ago
- Code for DUCK: Rumour Detection on Social Media by Modelling User and Comment Propagation Networks NAACL2022(https://aclanthology.org/202…☆23Jul 18, 2022Updated 4 years ago
- A Difficulty-Calibrated Benchmark for Building Terminal Agents☆30Feb 20, 2026Updated 6 months ago