A benchmark for evaluating LLM × harness performance.
☆101Aug 3, 2026Updated 3 weeks ago
Alternatives and similar repositories for PawBench
Users that are interested in PawBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A platform for running AI agents on realtime sport data and make predictions about game outcomes.☆48Jul 27, 2026Updated 3 weeks ago
- OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards☆799Aug 3, 2026Updated 3 weeks ago
- FlowLLM: Build LLM applications with ease.☆34Jun 27, 2026Updated last month
- A Multi-Agent Approach Integrating Socratic Guidance for Automated Prompt Optimization☆18Dec 15, 2025Updated 8 months ago
- ☆24Dec 15, 2025Updated 8 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- An in-the-wild benchmark for AI agents in the production harness.☆515Aug 17, 2026Updated last week
- [EMNLP 2026 Industry Track] EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions☆47Jun 26, 2026Updated 2 months ago
- ☆26May 29, 2026Updated 2 months ago
- Harness for running and evaluating AI agents against RL environments☆244Updated this week
- Reinforcing LLM Reasoning through Self-Training and Value-Guided Decoding☆18May 6, 2026Updated 3 months ago
- The implementation of <Factual Consistency Evaluation for Text Summarization via Counterfactual Estimation> in PyTorch.☆17Nov 11, 2021Updated 4 years ago
- ☆13Feb 22, 2024Updated 2 years ago
- UIUC CS 225 SP20 ZJUI Environment☆12May 30, 2020Updated 6 years ago
- Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement fine-tuning (RFT) of large language models (…☆693Aug 13, 2026Updated last week
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Fast and Slow Generating: An Empirical Study on Large and Small Language Models Collaborative Decoding.☆13Nov 19, 2024Updated last year
- A collection of ready-to-use Python sample agents built with AgentScope and AgentScope Runtime, covering use cases from CLI tools to full…☆340Apr 10, 2026Updated 4 months ago
- Dynamic dual-granularity skill bank for agentic RL, jointly evolving policy and skills to improve long-horizon decision making in agentic…☆69Apr 1, 2026Updated 4 months ago
- Agent-based implementation of RAG, incorporating AI agents into the RAG pipeline to orchestrate its components and perform additional act…☆20Feb 20, 2025Updated last year
- 一个强调工程化、可观测、可测试、可扩展的 RAG 项目。TraceRAG 的目标不是只把答案“生成出来”,而是把文档导入、切块、向量化、检索、带来源回答、评估与后续 tracing 拆成可独立验证的阶段,逐步演进成一个可维护、可解释、可复盘的生产级 RAG。☆15Apr 2, 2026Updated 4 months ago
- [NeurIPS 2025] Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models☆17Nov 2, 2025Updated 9 months ago
- Benchmark and execution environment for evaluating LLM agents on end-to-end AI Research. [ICLR 2026]☆39May 31, 2026Updated 2 months ago
- Code for the paper AdvST: Revisiting Data Augmentations for Single Domain Generalization (AAAI 2024)☆15May 6, 2024Updated 2 years ago
- A curated collection of skills around AgentScope ecosystem and CoPaw applications.☆123Apr 15, 2026Updated 4 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Academic papers and works related to SWE-bench and SWE-agents☆15Dec 8, 2025Updated 8 months ago
- Here is the repo for public scripts.☆12Jul 16, 2022Updated 4 years ago
- General Agent Benchmark for OpenClaw, made by Qwen Team, Alibaba Group.☆59Jun 10, 2026Updated 2 months ago
- [VLDB2026] Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards☆15Jul 8, 2026Updated last month
- A Library for intra-GPU/Inter-SM parallelsim☆12Aug 7, 2026Updated 2 weeks ago
- 【PyTorch】Easy-to-use package of Cognitive Evolutionary Search (CELS) for Click-Through Rate Prediction☆21Oct 29, 2023Updated 2 years ago
- ☆15Feb 26, 2026Updated 6 months ago
- [NAACL 2025] Source code for MMEvalPro, a more trustworthy and efficient benchmark for evaluating LMMs☆25Sep 26, 2024Updated last year
- [KDD'25] Combinatorial Optimization Perspective based Framework for Multi-behavior Recommendation☆20Jan 7, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- SkillsBench evaluates how well skills work and how effective agents are at using them.☆1,721Jul 23, 2026Updated last month
- ☆16Nov 17, 2024Updated last year
- MUA-RL: MULTI-TURN USER-INTERACTING AGENT REINFORCEMENT LEARNING FOR AGENTIC TOOL USE☆67Nov 5, 2025Updated 9 months ago
- Slowdown prediction module of Echo: Simulating Distributed Training at Scale☆13Jul 11, 2026Updated last month
- The code for paper "MemGym: a Long-Horizon Memory Environment for LLM Agents".☆21Jun 2, 2026Updated 2 months ago
- AutoLibra: Metric Induction for Agents from Open-Ended Human Feedback☆19Apr 23, 2026Updated 4 months ago
- The implementation of RAGSynth: Synthetic Data for Robust and Faithful RAG Component Optimization☆21May 26, 2025Updated last year