A benchmark for evaluating LLM × harness performance.
☆114Aug 3, 2026Updated 2 months ago
Alternatives and similar repositories for PawBench
Users that are interested in PawBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- LLM-powered MCP server for building financial deep-research agents, integrating web search, Crawl4AI scraping, and entity extraction into…☆26Feb 11, 2026Updated 7 months ago
- A platform for running AI agents on realtime sport data and make predictions about game outcomes.☆50Jul 27, 2026Updated 2 months ago
- OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards☆871Sep 11, 2026Updated 3 weeks ago
- FlowLLM: Build LLM applications with ease.☆36Jun 27, 2026Updated 3 months ago
- Multi-tenant fine-tuning for LLMs with Tinker-compatible API☆74Sep 29, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A Multi-Agent Approach Integrating Socratic Guidance for Automated Prompt Optimization☆18Dec 15, 2025Updated 9 months ago
- ☆21May 12, 2026Updated 4 months ago
- An in-the-wild benchmark for AI agents in the production harness.☆529Sep 18, 2026Updated 3 weeks ago
- [EMNLP 2026 Industry Track] EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions☆49Jun 26, 2026Updated 3 months ago
- Cutting-edge platform for LLM agent tuning. Deliver RL tuning with flexibility, reliability, speed, multi-agent optimization and realtime…☆240Aug 25, 2026Updated last month
- Harness for running and evaluating AI agents against RL environments☆283Updated this week
- Reinforcing LLM Reasoning through Self-Training and Value-Guided Decoding☆19May 6, 2026Updated 5 months ago
- Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement fine-tuning (RFT) of large language models (…☆706Oct 1, 2026Updated last week
- An Efficent BPE Algorithm Faster then Hugging Face Tokenizer's Implementation☆13Sep 30, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Fast and Slow Generating: An Empirical Study on Large and Small Language Models Collaborative Decoding.☆13Nov 19, 2024Updated last year
- A collection of ready-to-use Python sample agents built with AgentScope and AgentScope Runtime, covering use cases from CLI tools to full…☆352Apr 10, 2026Updated 6 months ago
- 一个强调工程化、可观测、可测试、可扩展的 RAG 项目。TraceRAG 的目标不是只把答案“生成出来”,而是把文档导入、切块、向量化、检索、带来源回答、评估与后续 tracing 拆成可独立验证的阶段,逐步演进成一个可维护、可解释、可复盘的生产级 RAG。☆15Apr 2, 2026Updated 6 months ago
- [NeurIPS 2025] Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models☆17Nov 2, 2025Updated 11 months ago
- Benchmark and execution environment for evaluating LLM agents on end-to-end AI Research. [ICLR 2026]☆42May 31, 2026Updated 4 months ago
- ☆30Mar 30, 2026Updated 6 months ago
- Academic papers and works related to SWE-bench and SWE-agents☆18Dec 8, 2025Updated 10 months ago
- Here is the repo for public scripts.☆12Jul 16, 2022Updated 4 years ago
- A curated collection of skills around AgentScope ecosystem and CoPaw applications.☆130Sep 10, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 【PyTorch】Easy-to-use package of Cognitive Evolutionary Search (CELS) for Click-Through Rate Prediction☆21Oct 29, 2023Updated 2 years ago
- General Agent Benchmark for OpenClaw, made by Qwen Team, Alibaba Group.☆61Jun 10, 2026Updated 4 months ago
- [NAACL 2025] Source code for MMEvalPro, a more trustworthy and efficient benchmark for evaluating LMMs☆25Sep 26, 2024Updated 2 years ago
- [KDD'25] Combinatorial Optimization Perspective based Framework for Multi-behavior Recommendation☆20Jan 7, 2025Updated last year
- Probabilistic data structures for processing very large datasets (MinHash, HyperLogLog)☆11Aug 20, 2015Updated 11 years ago
- SkillsBench evaluates how well skills work and how effective agents are at using them.☆1,836Jul 23, 2026Updated 2 months ago
- MUA-RL: MULTI-TURN USER-INTERACTING AGENT REINFORCEMENT LEARNING FOR AGENTIC TOOL USE☆68Nov 5, 2025Updated 11 months ago
- CAR-bench☆44Aug 27, 2026Updated last month
- Slowdown prediction module of Echo: Simulating Distributed Training at Scale☆13Jul 11, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆35May 16, 2025Updated last year
- The code for paper "MemGym: a Long-Horizon Memory Environment for LLM Agents".☆23Jun 2, 2026Updated 4 months ago
- 一些有趣的页面,使用 Github Pages 和 Vercel 部署☆15Feb 8, 2024Updated 2 years ago
- The implementation of RAGSynth: Synthetic Data for Robust and Faithful RAG Component Optimization☆21May 26, 2025Updated last year
- Doc Thinker: All-in-One RAG - document parsing, graph RAG, and evaluation☆32Sep 19, 2026Updated 3 weeks ago
- Code for "Practical Low-Rank Communication Compression in Decentralized Deep Learning"☆17Aug 4, 2020Updated 6 years ago
- [NeurIPS 2025] A multimodal agent that can interact with its own PC in a multimodal manner.☆39Sep 12, 2026Updated 3 weeks ago